I think this is evidence against the usefulness of standardized tests rather than support for ChatGPT
ChatGPT scores 80% on SAT reading/writing with collective chain of thought
71–80 of 126 posts
Re: ChatGPT scores 80% on SAT reading/writing with collective chain of thought
#72This is like Coca-Cola making the Covid-19 tests positive. Cola making the test positive doesn't mean that the test is fake or that the Cola has virus and ChatGPT scoring something on SAT doesn't mean it's equivalent to a person scoring the same and should study in Stanford. These tests are designed to pick people to limited number of spots and doesn't say much more than the persons' preparedness for the test.
Logically thinking the only way that ChatGPT could do well on a test of reasoning without engaging in reasoning is if the test is useless as a test of ability to reason.
Since humans use this test to select people with good reasoning skills, why do you think it is an inappropriate test for a robot?
Re: ChatGPT scores 80% on SAT reading/writing with collective chain of thought
#73This is like Coca-Cola making the Covid-19 tests positive. Cola making the test positive doesn't mean that the test is fake or that the Cola has virus and ChatGPT scoring something on SAT doesn't mean it's equivalent to a person scoring the same and should study in Stanford. These tests are designed to pick people to limited number of spots and doesn't say much more than the persons' preparedness for the test.
Peak denialism. I and thousands of others use GPT's general creative intelligence, problem-solving skills and ability to manipulate abstract thoughts daily. Logically thinking the only way that ChatGPT could do well on a test of reasoning without engaging in reasoning is if the test is useless as a test of ability to reason. Since humans use this test to select people with good reasoning skills, why do you think it i…
Sure, we can keep pretending this I guess.
Re: ChatGPT scores 80% on SAT reading/writing with collective chain of thought
#74Earlier quoted context omitted.
Yes, but ... what this shows is that testing for "memorization" isn't good enough. We all know people who have learned information and can give answers with confidence, even though they are just winging it, to put it politely. ChatGPT could be considered their equal on such tasks. So schools need to test that students are capable of working with the acquired information in a way that befits the level of the course, j…
The question is if we should expect these test to test on the black belt level or at a lower level belt? (i am not native English speaking so i don't know what addication means, I couldn't quite understand that sentence)
Re: ChatGPT scores 80% on SAT reading/writing with collective chain of thought
#75This is like Coca-Cola making the Covid-19 tests positive. Cola making the test positive doesn't mean that the test is fake or that the Cola has virus and ChatGPT scoring something on SAT doesn't mean it's equivalent to a person scoring the same and should study in Stanford. These tests are designed to pick people to limited number of spots and doesn't say much more than the persons' preparedness for the test.
Peak denialism. I and thousands of others use GPT's general creative intelligence, problem-solving skills and ability to manipulate abstract thoughts daily. Logically thinking the only way that ChatGPT could do well on a test of reasoning without engaging in reasoning is if the test is useless as a test of ability to reason. Since humans use this test to select people with good reasoning skills, why do you think it i…
ChatGPT has achieved 80% in this user's run, will refuse to play at all the next run, will get 45% the following run etc.
Stochastic randomisation makes the chatbot useful and provides these moments of impressive capability, it also means unlike humans that capability is in no way consistent.
Re: ChatGPT scores 80% on SAT reading/writing with collective chain of thought
#76This is like Coca-Cola making the Covid-19 tests positive. Cola making the test positive doesn't mean that the test is fake or that the Cola has virus and ChatGPT scoring something on SAT doesn't mean it's equivalent to a person scoring the same and should study in Stanford. These tests are designed to pick people to limited number of spots and doesn't say much more than the persons' preparedness for the test.
Peak denialism. I and thousands of others use GPT's general creative intelligence, problem-solving skills and ability to manipulate abstract thoughts daily. Logically thinking the only way that ChatGPT could do well on a test of reasoning without engaging in reasoning is if the test is useless as a test of ability to reason. Since humans use this test to select people with good reasoning skills, why do you think it i…
The test is not useless, the test's use case is to create a standardised datapoint to pick people to limited spots. That' what the test does, you need to attach false attributes to the test(like measures intelligence) to associate these attributes to the machine and this is a fallacy because although the test probably says something about the intelligence of the human, it does it by measuring meta attributes of the intelligence of the human(like speed of solving problems that the human was supposed to train on).
This is also why there is the "Goodhart's law" phenomenon where in this case students instead of learning physics might start putting all their effort on learning to solve physics tests and you end up with students with very high scores in physics but with no understanding of physics.
Re: ChatGPT scores 80% on SAT reading/writing with collective chain of thought
#77Earlier quoted context omitted.
Yes, but ... what this shows is that testing for "memorization" isn't good enough. We all know people who have learned information and can give answers with confidence, even though they are just winging it, to put it politely. ChatGPT could be considered their equal on such tasks. So schools need to test that students are capable of working with the acquired information in a way that befits the level of the course, j…
The question is if we should expect these test to test on the black belt level or at a lower level belt? (i am not native English speaking so i don't know what addication means, I couldn't quite understand that sentence)
Re: ChatGPT scores 80% on SAT reading/writing with collective chain of thought
#78Earlier quoted context omitted.
Peak denialism. I and thousands of others use GPT's general creative intelligence, problem-solving skills and ability to manipulate abstract thoughts daily. Logically thinking the only way that ChatGPT could do well on a test of reasoning without engaging in reasoning is if the test is useless as a test of ability to reason. Since humans use this test to select people with good reasoning skills, why do you think it i…
A human doing the same SATs will test roughly the same way every time. ChatGPT has achieved 80% in this user's run, will refuse to play at all the next run, will get 45% the following run etc. Stochastic randomisation makes the chatbot useful and provides these moments of impressive capability, it also means unlike humans that capability is in no way consistent.
> ChatGPT has achieved 80% in this user's run, will refuse to play at all the next run, will get 45% the following run etc.
While I agree with your conclusion, I want to slightly modify the first sentence. I know this is just one anecdote but I took the SAT 1 twice and did somewhat better the second time. I think this is common as you are more familiar with the test the second time.
Re: ChatGPT scores 80% on SAT reading/writing with collective chain of thought
#79Earlier quoted context omitted.
I just tested it with clearer language... and the response is incorrect: Question: There are 5 dogs in a room right now. The previous day an additional dog was purchased and added to the room. How many dogs are in the room right now? Answer: There are 6 dogs in the room right now. The statement says that the previous day an additional dog was purchased and added to the room, so the total number of dogs in the room ri…
I think the word "additional" is pretty confusing here, because it begs the question of "additional to what?" (ChatGPT still doesn't get it right if you remove "additional")
Re: ChatGPT scores 80% on SAT reading/writing with collective chain of thought
#80Earlier quoted context omitted.
HN seems to hate both, and technology in general.
Being critical doesn't necessarily equal hate. It's not surprising a forum for coders would have people criticize and not just rave about tech like people who do not understand how it works and to whom it's basically magic.