Instead, we are comparing the understanding of a person (who takes the SAT and scores something) against the logical soundness of a set of expressions written by other people (the training data), all weighted by GPT's ability to semantically transform those expressions into new expressions (answers).
ChatGPT scores 80% on SAT reading/writing with collective chain of thought
91–100 of 126 posts
Re: ChatGPT scores 80% on SAT reading/writing with collective chain of thought
#92This comes up a lot in these prompt engineering examples: Why is ChatGPT so stupid and evasive by default, speaking in dumb generalities, but gets very smart when you say "pretend that you are a smart person"?
Re: ChatGPT scores 80% on SAT reading/writing with collective chain of thought
#93This is like Coca-Cola making the Covid-19 tests positive. Cola making the test positive doesn't mean that the test is fake or that the Cola has virus and ChatGPT scoring something on SAT doesn't mean it's equivalent to a person scoring the same and should study in Stanford. These tests are designed to pick people to limited number of spots and doesn't say much more than the persons' preparedness for the test.
Your body was able to transform what was in the soda (sugary water) into a completely different thing (ammonia and water), while preserving a key attribute (water).
Our assertion that the preserved attribute is "key" is the act of us choosing what to be impressed about.
GPT was trained to be impressive by preserving the logical soundness of expression when semantically transforming expression. It did not explicitly choose the logical conclusions: only the semantic match that contains them.
Expressing logical conclusions in natural language is inherently ambiguous, which makes this a very large and incoherent problem space.
Re: ChatGPT scores 80% on SAT reading/writing with collective chain of thought
#94This is like Coca-Cola making the Covid-19 tests positive. Cola making the test positive doesn't mean that the test is fake or that the Cola has virus and ChatGPT scoring something on SAT doesn't mean it's equivalent to a person scoring the same and should study in Stanford. These tests are designed to pick people to limited number of spots and doesn't say much more than the persons' preparedness for the test.
Peak denialism. I and thousands of others use GPT's general creative intelligence, problem-solving skills and ability to manipulate abstract thoughts daily. Logically thinking the only way that ChatGPT could do well on a test of reasoning without engaging in reasoning is if the test is useless as a test of ability to reason. Since humans use this test to select people with good reasoning skills, why do you think it i…
That's a bit condescending.
ChatGPT isn't remotely close to "general creative intelligence, problem-solving skills and ability to manipulate abstract thoughts daily". It only pretends, in a pretty impressive way. A bit like you can train your dog to play dead when you point your finger at them and say "bang", but you should not conclude from it that your dog understands the concept of death.
> Since humans use this test to select people with good reasoning skills, why do you think it is an inappropriate test for a robot?
Well it is already not a perfect solution for humans. When you use a metric to rate people, then they start optimizing for this metric. Typical example is academic rankings: if you are ranked based on how many papers you publish, you will publish a lot. Doesn't mean you will do good research.
Now even if SAT was a good way to rank humans, that does not say at all it would be a good way for robots. If you test whether a cat is stupid or not by asking it to drive a car, then obviously it will be stupid. Asking ChatGPT to pas the SAT tests is probably like asking a student to pass it with infinite time and an internet connection. I guess that's not how SAT exams are usually done, right? (I am not in the US).
Re: ChatGPT scores 80% on SAT reading/writing with collective chain of thought
#95Earlier quoted context omitted.
Yeah this seems like a trick question that can be interpreted multiple ways, especially if english isn't your native language. This is exactly why I prefer the precision of code over natural language for inputs (and why Google's botsplaining of "what you really meant to search" instead of just searching exactly the query as-is is so frustrating). In a test, this seems like the kind of question they use to justify not…
I just tested it with clearer language... and the response is incorrect: Question: There are 5 dogs in a room right now. The previous day an additional dog was purchased and added to the room. How many dogs are in the room right now? Answer: There are 6 dogs in the room right now. The statement says that the previous day an additional dog was purchased and added to the room, so the total number of dogs in the room ri…
Re: ChatGPT scores 80% on SAT reading/writing with collective chain of thought
#96Earlier quoted context omitted.
There is one argument against this though that is somehow brewing in my brain. The end goal is to internalize the subject so that you can extrapolate your own insights. "This happened in WWII so if these conditions starts to appear it risks to happen again". Just to take history as an example. All other subjects have similar examples. And to reach that kind of insight you need to start filling the brain with informat…
Yes, but ... what this shows is that testing for "memorization" isn't good enough. We all know people who have learned information and can give answers with confidence, even though they are just winging it, to put it politely. ChatGPT could be considered their equal on such tasks. So schools need to test that students are capable of working with the acquired information in a way that befits the level of the course, j…
Now it would be wonderful if we could just motivate students to understand and learn the material, without putting so much pressure by testing them all the time.
Re: ChatGPT scores 80% on SAT reading/writing with collective chain of thought
#97Earlier quoted context omitted.
I interpreted it as a linear chain of events, end result being 28.
If you interpret the room with chairs in it as a storage room, all of the facts are relevant: - we had 19 chairs in the storage room - we used 10 chairs (removing them from storage) - and then we purchased 9 more chairs (adding them to storage) The answer from chatGPT is consistent with English and makes use of all the facts; either the question is testing for deception (intentionally sharing misleading information)…
Possible answers:
- 19 chairs, least likely answer given the original
- 18 chairs, 10 of original 19 removed and 9 added
- 28 chairs, 19 chairs in the room with 10 in active use and 9 new chairs added
Just seems like a question a human would fail too. Probably opens a can of worms on interview questions and personal bias etc too in a more general sense. Not bias in the political sense, just how interview questions encode a lot of the asker's assumptions.
Edit: the only unambiguous phrasing I can think of is "19 chairs are in a room, 9 of the chairs are green, how many chairs are in the room?".
Re: ChatGPT scores 80% on SAT reading/writing with collective chain of thought
#98Earlier quoted context omitted.
| Q: The room has 19 chairs. We are using 10 and then bought 9 more, how many chairs do we have in the room? I don't know if you made this question up, but I think this question has issues (I as a reader find it meaningless). You say you bought 9 more chairs, but you don't say you put those 9 chairs into the room. Also you don't specify "we" own any of the 19 chairs in the room initially, but at the end you ask how m…
The room has 19 chairs so the only sensible interpretation of the second sentence is as an explanation for why there are 19 chairs. That is the part ChatGPT missed, as it does not try to interpret ambiguous things that way.
Re: ChatGPT scores 80% on SAT reading/writing with collective chain of thought
#99Earlier quoted context omitted.
No, the answer is that it's impossible to know with the information provided. We don't know if "used" implies removal, nor if "bought" implies addition. You could make reasonable arguments in any direction.
Your post has 33 words. You used 20 words and then typed 13 more, how many words does your post have?
Re: ChatGPT scores 80% on SAT reading/writing with collective chain of thought
#100Earlier quoted context omitted.
If you interpret the room with chairs in it as a storage room, all of the facts are relevant: - we had 19 chairs in the storage room - we used 10 chairs (removing them from storage) - and then we purchased 9 more chairs (adding them to storage) The answer from chatGPT is consistent with English and makes use of all the facts; either the question is testing for deception (intentionally sharing misleading information)…
This is a very tortuous way to try to make this question ambiguous. Read the words in the problem and stop inserting additional facts: The room has 19 chairs. We are using 10 and then bought 9 more, how many chairs do we have in the room? The room is the room (what type of room is not relevant). There are 19 chairs in it. The 19 chairs most likely got there when 9 more chairs were purchased after 10 were used, but th…