Live data from Hacker News

ChatGPT scores 80% on SAT reading/writing with collective chain of thought

old.reddit.com

21–30 of 126 posts

Re: ChatGPT scores 80% on SAT reading/writing with collective chain of thought

#21
post #14

Using chain of thought [slightly modified] example from https://arxiv.org/abs/2201.11903 (ref'd in OP) Prompt: Q: Roger has 5 tennis balls. He buys 2 more cans of tennis balls. Each can has 3 tennis balls. How many tennis balls does he have now? A: Roger started with 5 balls. 2 cans of 3 tennis balls each is 6 tennis balls. 5 + 6 = 11. The answer is 11. Q: The room has 19 chairs. We are using 10 and then bought 9 mor…

| Q: The room has 19 chairs. We are using 10 and then bought 9 more, how many chairs do we have in the room? I don't know if you made this question up, but I think this question has issues (I as a reader find it meaningless). You say you bought 9 more chairs, but you don't say you put those 9 chairs into the room. Also you don't specify "we" own any of the 19 chairs in the room initially, but at the end you ask how m…

None of that matters. The room has 19 chairs. The answer is 19.

Re: ChatGPT scores 80% on SAT reading/writing with collective chain of thought

#22

I think this is evidence against the usefulness of standardized tests rather than support for ChatGPT

What I find more concerning is the 100% scores that ChatGPT has been able to achieve on AP CSA and APUSH, among others.

It makes sense to teach 3rd graders 2+2, but if college board is creating "college level courses" with problems that can be easily solved through whats still a fairly new technology, we need to reevaluate what we're teaching in schools.

So many AP courses are just memorization and identifying formulas. I'd bet the majority of AP CSA students wouldn't fare very well when tasked with detailing how they'd approach a project, whereas kids in IB schools etc. would find it straightforward.

Re: ChatGPT scores 80% on SAT reading/writing with collective chain of thought

#24
post #10

Earlier quoted context omitted.

And answer from the new version: > After using 10 chairs, we have 19 chairs - 10 chairs = 9 chairs in the room. Then after buying 9 more chairs, we have 9 chairs + 9 chairs = 18 chairs in the room.

Which version are you using? I'm on " ChatGPT Jan 9 Version ", and not only it got right, but also why someone would (incorrectly) guess 18. See edit above. But of course, there's an initial entropy to the models, so it's very possible that depending on the initial state it converged into the wrong (but yet probabilistically plausible) answer.

English is not a native language for me, and I'm really not sure what is a correct answer.

The question is "how many chairs do we have in a room". But those 19 initial chairs was stated without any reference to "we", they are just chairs in the room. "We" are using 9 chairs, so we in some sense claim ownership on these. And "we" bought 9 more chairs, so they are also owned by "we". 9+9 is 18, and 10 chars that in the room but "we" do not have them.

Think of a common room shared by two not exactly friendly groups of people. Each owns some chairs and is ready to defend its rights. So if "we" is one of groups and we need more chairs the only peaceful way to get them is to buy more chairs.

I believe that this understanding of the question seems correct to me because I do not see some implicit assumptions, and it is due to my English is not good enough (I just know, I cannot reliably solve logic puzzles written in English). Can you please explain on this example why my reasoning is wrong?

Re: ChatGPT scores 80% on SAT reading/writing with collective chain of thought

#25
post #21
post #14

Earlier quoted context omitted.

| Q: The room has 19 chairs. We are using 10 and then bought 9 more, how many chairs do we have in the room? I don't know if you made this question up, but I think this question has issues (I as a reader find it meaningless). You say you bought 9 more chairs, but you don't say you put those 9 chairs into the room. Also you don't specify "we" own any of the 19 chairs in the room initially, but at the end you ask how m…

None of that matters. The room has 19 chairs. The answer is 19.

I interpreted it as a linear chain of events, end result being 28.

Re: ChatGPT scores 80% on SAT reading/writing with collective chain of thought

#26

I think this is evidence against the usefulness of standardized tests rather than support for ChatGPT

What I find more concerning is the 100% scores that ChatGPT has been able to achieve on AP CSA and APUSH, among others. It makes sense to teach 3rd graders 2+2, but if college board is creating "college level courses" with problems that can be easily solved through whats still a fairly new technology, we need to reevaluate what we're teaching in schools. So many AP courses are just memorization and identifying formul…

I don't find it at all concerning that e.g. a history course is primarily tested and easily passed by memorization. There's nothing wrong with teaching context instead of pure skills all the time IMO. I do agree it'd be good to see deeper computer science as an early option in the US but at the same time I'm not sure whether or not something like ChatGPT can pass is really a good indicator for why.

Re: ChatGPT scores 80% on SAT reading/writing with collective chain of thought

#27

I think this is evidence against the usefulness of standardized tests rather than support for ChatGPT

What I find more concerning is the 100% scores that ChatGPT has been able to achieve on AP CSA and APUSH, among others. It makes sense to teach 3rd graders 2+2, but if college board is creating "college level courses" with problems that can be easily solved through whats still a fairly new technology, we need to reevaluate what we're teaching in schools. So many AP courses are just memorization and identifying formul…

The AP CSA pass surprises me (I've never taken it myself). I had this thought because I tried to use chatgpt for work programming stuff and it got stuff wrong very frequently. So frequently, and to a great enough degree, that I had to keep re-circumscribing what kind of problems were worth my time to throw at it, and now it's a tiny cross-section.

Re: ChatGPT scores 80% on SAT reading/writing with collective chain of thought

#28
post #21
post #14

Earlier quoted context omitted.

| Q: The room has 19 chairs. We are using 10 and then bought 9 more, how many chairs do we have in the room? I don't know if you made this question up, but I think this question has issues (I as a reader find it meaningless). You say you bought 9 more chairs, but you don't say you put those 9 chairs into the room. Also you don't specify "we" own any of the 19 chairs in the room initially, but at the end you ask how m…

None of that matters. The room has 19 chairs. The answer is 19.

No, the answer is that it's impossible to know with the information provided. We don't know if "used" implies removal, nor if "bought" implies addition. You could make reasonable arguments in any direction.

Re: ChatGPT scores 80% on SAT reading/writing with collective chain of thought

#29
post #21

Earlier quoted context omitted.

None of that matters. The room has 19 chairs. The answer is 19.

I interpreted it as a linear chain of events, end result being 28.

Yeah this seems like a trick question that can be interpreted multiple ways, especially if english isn't your native language.

This is exactly why I prefer the precision of code over natural language for inputs (and why Google's botsplaining of "what you really meant to search" instead of just searching exactly the query as-is is so frustrating).

In a test, this seems like the kind of question they use to justify not giving a perfect grade because whatever answer you pick, they can claim it was the other one because both interpretations are reasonable.

Re: ChatGPT scores 80% on SAT reading/writing with collective chain of thought

#30
post #21

Earlier quoted context omitted.

None of that matters. The room has 19 chairs. The answer is 19.

I interpreted it as a linear chain of events, end result being 28.

If you interpret the room with chairs in it as a storage room, all of the facts are relevant:

- we had 19 chairs in the storage room

- we used 10 chairs (removing them from storage)

- and then we purchased 9 more chairs (adding them to storage)

The answer from chatGPT is consistent with English and makes use of all the facts; either the question is testing for deception (intentionally sharing misleading information) or chatGPT made the correct inference, by assuming good faith and using all the provided information.

This question says something about us — that we’d assume deception ahead of unclear reference to a storage room by a good faith speaker.

Post reply on HN