Live data from Hacker News

ChatGPT scores 80% on SAT reading/writing with collective chain of thought

old.reddit.com

31–40 of 126 posts

Re: ChatGPT scores 80% on SAT reading/writing with collective chain of thought

#31
post #28
post #21

Earlier quoted context omitted.

None of that matters. The room has 19 chairs. The answer is 19.

No, the answer is that it's impossible to know with the information provided. We don't know if "used" implies removal, nor if "bought" implies addition. You could make reasonable arguments in any direction.

Homo sapiens are safe. Lawyers are not.

Re: ChatGPT scores 80% on SAT reading/writing with collective chain of thought

#32
post #24

Earlier quoted context omitted.

Which version are you using? I'm on " ChatGPT Jan 9 Version ", and not only it got right, but also why someone would (incorrectly) guess 18. See edit above. But of course, there's an initial entropy to the models, so it's very possible that depending on the initial state it converged into the wrong (but yet probabilistically plausible) answer.

English is not a native language for me, and I'm really not sure what is a correct answer. The question is "how many chairs do we have in a room". But those 19 initial chairs was stated without any reference to "we", they are just chairs in the room. "We" are using 9 chairs, so we in some sense claim ownership on these. And "we" bought 9 more chairs, so they are also owned by "we". 9+9 is 18, and 10 chars that in the…

You're overthinking. There are no hidden tricks. It's simple logic: the room has 19 chairs. The number of chairs in use is irrelevant. You got 9 more chairs; how many chairs you have in this room now? Of course the answer is 28.

It's a problem created to test the limits of interpretation of large language models. Here's another one: "What would be the gender of the first female president of the United States?" [1]

In previous versions it used to get this wrong, but the current version of ChatGPT nails it every time. Seems it's getting better, likely incorporating human feedback.

[1] https://www.nytimes.com/2023/01/06/podcasts/transcript-ezra-...

Re: ChatGPT scores 80% on SAT reading/writing with collective chain of thought

#33

Using chain of thought [slightly modified] example from https://arxiv.org/abs/2201.11903 (ref'd in OP) Prompt: Q: Roger has 5 tennis balls. He buys 2 more cans of tennis balls. Each can has 3 tennis balls. How many tennis balls does he have now? A: Roger started with 5 balls. 2 cans of 3 tennis balls each is 6 tennis balls. 5 + 6 = 11. The answer is 11. Q: The room has 19 chairs. We are using 10 and then bought 9 mor…

> Q: The room has 19 chairs. We are using 10 and then bought 9 more, how many chairs do we have in the room?

This is a terrible question and i'm not surprised chatGPT is subtly telling you it can't figure it out.

"The room has 19 chairs" 19 chairs belong to the room or 19 chairs are in the room?

"We are using 10" are "we" in the room too? are we using them for sitting, or does "used" imply consumption e.g. used for firewood? is this statement about the room's chairs? is this statement about chairs at all?

"and then bought 9 more," presumably we left the room to buy the chairs, did we bring them back into the room? if we didn't leave the room to buy, i.e bought them online, have they arrived yet? it's still not clear if we are in or have ever been in the room.

> how many chairs do we have in the room?

how many chairs do _we have_? does the room's initial set of chairs belong to us? if not, do we just count the ones we bought?

Re: ChatGPT scores 80% on SAT reading/writing with collective chain of thought

#34
post #24

Earlier quoted context omitted.

English is not a native language for me, and I'm really not sure what is a correct answer. The question is "how many chairs do we have in a room". But those 19 initial chairs was stated without any reference to "we", they are just chairs in the room. "We" are using 9 chairs, so we in some sense claim ownership on these. And "we" bought 9 more chairs, so they are also owned by "we". 9+9 is 18, and 10 chars that in the…

You're overthinking. There are no hidden tricks. It's simple logic: the room has 19 chairs. The number of chairs in use is irrelevant. You got 9 more chairs; how many chairs you have in this room now? Of course the answer is 28. It's a problem created to test the limits of interpretation of large language models. Here's another one: " What would be the gender of the first female president of the United States? " [1]…

> You're overthinking. There are no hidden tricks. It's simple logic: the room has 19 chairs.

That's a generous interpretation. I have read questions in textbooks that have made far worse logical errors, leaving the reader to infer one of multiple possible meanings.

Re: ChatGPT scores 80% on SAT reading/writing with collective chain of thought

#35

Earlier quoted context omitted.

I interpreted it as a linear chain of events, end result being 28.

Yeah this seems like a trick question that can be interpreted multiple ways, especially if english isn't your native language. This is exactly why I prefer the precision of code over natural language for inputs (and why Google's botsplaining of "what you really meant to search" instead of just searching exactly the query as-is is so frustrating). In a test, this seems like the kind of question they use to justify not…

I just tested it with clearer language... and the response is incorrect:

Question: There are 5 dogs in a room right now. The previous day an additional dog was purchased and added to the room. How many dogs are in the room right now?

Answer: There are 6 dogs in the room right now. The statement says that the previous day an additional dog was purchased and added to the room, so the total number of dogs in the room right now is 5 + 1 = 6.

Edit: Ok a comment below pointed out that the question is still unclear. See below for the fix and the response. It is still incorrect:

Q: There are 5 dogs in a room on 1 January 2023. The previous day an additional dog was purchased and added to the room. How many dogs are in the room on 1 January 2023?

A: There are 6 dogs in the room on 1 January 2023. The previous day, one additional dog was purchased and added to the room, which means there are 5 dogs + 1 additional dog = 6 dogs in the room on 1 January 2023.

Re: ChatGPT scores 80% on SAT reading/writing with collective chain of thought

#36

I think this is evidence against the usefulness of standardized tests rather than support for ChatGPT

What I find more concerning is the 100% scores that ChatGPT has been able to achieve on AP CSA and APUSH, among others. It makes sense to teach 3rd graders 2+2, but if college board is creating "college level courses" with problems that can be easily solved through whats still a fairly new technology, we need to reevaluate what we're teaching in schools. So many AP courses are just memorization and identifying formul…

There is one argument against this though that is somehow brewing in my brain.

The end goal is to internalize the subject so that you can extrapolate your own insights.

"This happened in WWII so if these conditions starts to appear it risks to happen again". Just to take history as an example. All other subjects have similar examples.

And to reach that kind of insight you need to start filling the brain with information. That may be called memorization. The memorization is just one path of the road towards the black belt. If you skip it you will be lost.

Re: ChatGPT scores 80% on SAT reading/writing with collective chain of thought

#37

Earlier quoted context omitted.

Yeah this seems like a trick question that can be interpreted multiple ways, especially if english isn't your native language. This is exactly why I prefer the precision of code over natural language for inputs (and why Google's botsplaining of "what you really meant to search" instead of just searching exactly the query as-is is so frustrating). In a test, this seems like the kind of question they use to justify not…

I just tested it with clearer language... and the response is incorrect: Question: There are 5 dogs in a room right now. The previous day an additional dog was purchased and added to the room. How many dogs are in the room right now? Answer: There are 6 dogs in the room right now. The statement says that the previous day an additional dog was purchased and added to the room, so the total number of dogs in the room ri…

Reading the question, I thought the answer was six as well. Such a question is so misleading that I would hope that an AI answers as it has in this case, because not all problems are fully logically specified, and you need to remain robust against the question being posed incorrectly.

In other words, I would posit the probability of you having meant for the answer to be 5 low, because the question itself becomes trivial. For a system that needs to deal with many people asking questions, this robustness is helpful in my view.

In other words, it's not totally clear that time has not passed between sentences in your prompt. So the second instance of "now" could be a different "now".

Re: ChatGPT scores 80% on SAT reading/writing with collective chain of thought

#38
post #37

Earlier quoted context omitted.

I just tested it with clearer language... and the response is incorrect: Question: There are 5 dogs in a room right now. The previous day an additional dog was purchased and added to the room. How many dogs are in the room right now? Answer: There are 6 dogs in the room right now. The statement says that the previous day an additional dog was purchased and added to the room, so the total number of dogs in the room ri…

Reading the question, I thought the answer was six as well. Such a question is so misleading that I would hope that an AI answers as it has in this case, because not all problems are fully logically specified, and you need to remain robust against the question being posed incorrectly. In other words, I would posit the probability of you having meant for the answer to be 5 low, because the question itself becomes triv…

Good point, I added a clearer version of the question to my comment above. It still has trouble with it.

Re: ChatGPT scores 80% on SAT reading/writing with collective chain of thought

#39
post #37

Earlier quoted context omitted.

I just tested it with clearer language... and the response is incorrect: Question: There are 5 dogs in a room right now. The previous day an additional dog was purchased and added to the room. How many dogs are in the room right now? Answer: There are 6 dogs in the room right now. The statement says that the previous day an additional dog was purchased and added to the room, so the total number of dogs in the room ri…

Reading the question, I thought the answer was six as well. Such a question is so misleading that I would hope that an AI answers as it has in this case, because not all problems are fully logically specified, and you need to remain robust against the question being posed incorrectly. In other words, I would posit the probability of you having meant for the answer to be 5 low, because the question itself becomes triv…

Reading and understanding exactly what they are asking for is a part of the test, adding superfluous information to problems is perfectly valid to test that students actually understand instead of trying to pattern match to solutions they memorized.

Re: ChatGPT scores 80% on SAT reading/writing with collective chain of thought

#40
post #10

Earlier quoted context omitted.

And answer from the new version: > After using 10 chairs, we have 19 chairs - 10 chairs = 9 chairs in the room. Then after buying 9 more chairs, we have 9 chairs + 9 chairs = 18 chairs in the room.

Which version are you using? I'm on " ChatGPT Jan 9 Version ", and not only it got right, but also why someone would (incorrectly) guess 18. See edit above. But of course, there's an initial entropy to the models, so it's very possible that depending on the initial state it converged into the wrong (but yet probabilistically plausible) answer.

I had additional context in that chat session and whole/part of the whole session is used as input for every output. I think it was fine-tuned on that example from the paper, that's why it can answer it in isolation.

This is interesting, I asked again with different irrelevant to math context in my previous inputs and here is the answer:

> You have 28 chairs in the room. If you had 10 chairs and then bought 9 more, you would have a total of 10 + 9 = 28 chairs in the room.

10 + 9 = 28

It's like you say. Statistics of words change based on the context, so you will get different (correct or incorrect) answer depending on the context of your previous inputs.

Post reply on HN