Live data from Hacker News

ChatGPT scores 80% on SAT reading/writing with collective chain of thought

old.reddit.com

81–90 of 126 posts

Re: ChatGPT scores 80% on SAT reading/writing with collective chain of thought

#81
post #65

Earlier quoted context omitted.

HN seems to hate both, and technology in general.

Being critical doesn't necessarily equal hate. It's not surprising a forum for coders would have people criticize and not just rave about tech like people who do not understand how it works and to whom it's basically magic.

All I see is unsubstantial and dismissive snark.

Re: ChatGPT scores 80% on SAT reading/writing with collective chain of thought

#82
post #28
post #21

Earlier quoted context omitted.

None of that matters. The room has 19 chairs. The answer is 19.

No, the answer is that it's impossible to know with the information provided. We don't know if "used" implies removal, nor if "bought" implies addition. You could make reasonable arguments in any direction.

Your post has 33 words. You used 20 words and then typed 13 more, how many words does your post have?

Re: ChatGPT scores 80% on SAT reading/writing with collective chain of thought

#83

Earlier quoted context omitted.

I interpreted it as a linear chain of events, end result being 28.

If you interpret the room with chairs in it as a storage room, all of the facts are relevant: - we had 19 chairs in the storage room - we used 10 chairs (removing them from storage) - and then we purchased 9 more chairs (adding them to storage) The answer from chatGPT is consistent with English and makes use of all the facts; either the question is testing for deception (intentionally sharing misleading information)…

This is a very tortuous way to try to make this question ambiguous. Read the words in the problem and stop inserting additional facts: The room has 19 chairs. We are using 10 and then bought 9 more, how many chairs do we have in the room?

The room is the room (what type of room is not relevant). There are 19 chairs in it. The 19 chairs most likely got there when 9 more chairs were purchased after 10 were used, but that is not relevant to the question: how many chairs are in the room? 19.

A := 19

B := 10

C := 9

What does A equal?

Re: ChatGPT scores 80% on SAT reading/writing with collective chain of thought

#85
post #75

Earlier quoted context omitted.

Peak denialism. I and thousands of others use GPT's general creative intelligence, problem-solving skills and ability to manipulate abstract thoughts daily. Logically thinking the only way that ChatGPT could do well on a test of reasoning without engaging in reasoning is if the test is useless as a test of ability to reason. Since humans use this test to select people with good reasoning skills, why do you think it i…

A human doing the same SATs will test roughly the same way every time. ChatGPT has achieved 80% in this user's run, will refuse to play at all the next run, will get 45% the following run etc. Stochastic randomisation makes the chatbot useful and provides these moments of impressive capability, it also means unlike humans that capability is in no way consistent.

ChatGPT isn't a human, it's a mix of data sampled from many humans. Random humans are often bad at SAT.

Re: ChatGPT scores 80% on SAT reading/writing with collective chain of thought

#86

I think this is evidence against the usefulness of standardized tests rather than support for ChatGPT

There is a leap in this logic- going from the fact that SAT is solvable by a LLM to the conclusion that the SAT is useless. Questions about addition can be solved using calculators, does that mean that teaching addition is useless? Also there is an underlying assumption that memorization is a useless skill, which I'd argue against, as in today's world most innovation is done by identifying patterns between discipline…

Well it shows that you can do well on SAT with basic word association without understanding and explaining the ideas of the passage.

The "Joe Bloggs" theory has known this for ages.

Re: ChatGPT scores 80% on SAT reading/writing with collective chain of thought

#87
post #68

This is like Coca-Cola making the Covid-19 tests positive. Cola making the test positive doesn't mean that the test is fake or that the Cola has virus and ChatGPT scoring something on SAT doesn't mean it's equivalent to a person scoring the same and should study in Stanford. These tests are designed to pick people to limited number of spots and doesn't say much more than the persons' preparedness for the test.

There’s a reason the SATs are not used in most of the world. They select for test takers. Kind of like the problem with the big tech style interview.

A human and a robot are both good at taking a bad test. This doesn’t say much about the robot.

Re: ChatGPT scores 80% on SAT reading/writing with collective chain of thought

#88
post #67

Earlier quoted context omitted.

There is a leap in this logic- going from the fact that SAT is solvable by a LLM to the conclusion that the SAT is useless. Questions about addition can be solved using calculators, does that mean that teaching addition is useless? Also there is an underlying assumption that memorization is a useless skill, which I'd argue against, as in today's world most innovation is done by identifying patterns between discipline…

not so much memorization as decomposition of numerous examples into logic based knowledge. As a simple example, knowing volumes of words is not as useful as knowing etymology since etymology brings you a deeper structure.

I'd say that's subjective, depending on context and individuals. Sure etymology might be amazing to know for a lit major or author, but for someone who wants to just be able to write clearly, knowing a lot of words is useful.

Re: ChatGPT scores 80% on SAT reading/writing with collective chain of thought

#89
post #86

Earlier quoted context omitted.

There is a leap in this logic- going from the fact that SAT is solvable by a LLM to the conclusion that the SAT is useless. Questions about addition can be solved using calculators, does that mean that teaching addition is useless? Also there is an underlying assumption that memorization is a useless skill, which I'd argue against, as in today's world most innovation is done by identifying patterns between discipline…

Well it shows that you can do well on SAT with basic word association without understanding and explaining the ideas of the passage. The "Joe Bloggs" theory has known this for ages.

Yeah but that also says nothing about how useful or useless that is

Re: ChatGPT scores 80% on SAT reading/writing with collective chain of thought

#90
post #85
post #75

Earlier quoted context omitted.

A human doing the same SATs will test roughly the same way every time. ChatGPT has achieved 80% in this user's run, will refuse to play at all the next run, will get 45% the following run etc. Stochastic randomisation makes the chatbot useful and provides these moments of impressive capability, it also means unlike humans that capability is in no way consistent.

ChatGPT isn't a human, it's a mix of data sampled from many humans. Random humans are often bad at SAT.

It isn't trained on the data from a random set of humans.

It's trained on the human data that was chosen as more impressive, and trained to pull the most impressive parts out of that data.

Fundamentally, while GPT may be "learned", it is not "logical". Instead of learning potential logical conclusions, it has learned potential semantic responses. It doesn't match the meaning of words to other meanings: it matchesa set of expressions to other expressions.

The impressive part is that the logical conclusions already present in the training data get presented by GPT in a semantically coherent way.

Post reply on HN