Earlier quoted context omitted.
HN seems to hate both, and technology in general.
Being critical doesn't necessarily equal hate. It's not surprising a forum for coders would have people criticize and not just rave about tech like people who do not understand how it works and to whom it's basically magic.
ChatGPT scores 80% on SAT reading/writing with collective chain of thought
81–90 of 126 posts
Re: ChatGPT scores 80% on SAT reading/writing with collective chain of thought
#82Earlier quoted context omitted.
None of that matters. The room has 19 chairs. The answer is 19.
No, the answer is that it's impossible to know with the information provided. We don't know if "used" implies removal, nor if "bought" implies addition. You could make reasonable arguments in any direction.
Re: ChatGPT scores 80% on SAT reading/writing with collective chain of thought
#83Earlier quoted context omitted.
I interpreted it as a linear chain of events, end result being 28.
If you interpret the room with chairs in it as a storage room, all of the facts are relevant: - we had 19 chairs in the storage room - we used 10 chairs (removing them from storage) - and then we purchased 9 more chairs (adding them to storage) The answer from chatGPT is consistent with English and makes use of all the facts; either the question is testing for deception (intentionally sharing misleading information)…
The room is the room (what type of room is not relevant). There are 19 chairs in it. The 19 chairs most likely got there when 9 more chairs were purchased after 10 were used, but that is not relevant to the question: how many chairs are in the room? 19.
A := 19
B := 10
C := 9
What does A equal?
Re: ChatGPT scores 80% on SAT reading/writing with collective chain of thought
#84Why is ChatGPT so stupid and evasive by default, speaking in dumb generalities, but gets very smart when you say "pretend that you are a smart person"?
Re: ChatGPT scores 80% on SAT reading/writing with collective chain of thought
#85Earlier quoted context omitted.
Peak denialism. I and thousands of others use GPT's general creative intelligence, problem-solving skills and ability to manipulate abstract thoughts daily. Logically thinking the only way that ChatGPT could do well on a test of reasoning without engaging in reasoning is if the test is useless as a test of ability to reason. Since humans use this test to select people with good reasoning skills, why do you think it i…
A human doing the same SATs will test roughly the same way every time. ChatGPT has achieved 80% in this user's run, will refuse to play at all the next run, will get 45% the following run etc. Stochastic randomisation makes the chatbot useful and provides these moments of impressive capability, it also means unlike humans that capability is in no way consistent.
Re: ChatGPT scores 80% on SAT reading/writing with collective chain of thought
#86I think this is evidence against the usefulness of standardized tests rather than support for ChatGPT
There is a leap in this logic- going from the fact that SAT is solvable by a LLM to the conclusion that the SAT is useless. Questions about addition can be solved using calculators, does that mean that teaching addition is useless? Also there is an underlying assumption that memorization is a useless skill, which I'd argue against, as in today's world most innovation is done by identifying patterns between discipline…
The "Joe Bloggs" theory has known this for ages.
Re: ChatGPT scores 80% on SAT reading/writing with collective chain of thought
#87This is like Coca-Cola making the Covid-19 tests positive. Cola making the test positive doesn't mean that the test is fake or that the Cola has virus and ChatGPT scoring something on SAT doesn't mean it's equivalent to a person scoring the same and should study in Stanford. These tests are designed to pick people to limited number of spots and doesn't say much more than the persons' preparedness for the test.
A human and a robot are both good at taking a bad test. This doesn’t say much about the robot.
Re: ChatGPT scores 80% on SAT reading/writing with collective chain of thought
#88Earlier quoted context omitted.
There is a leap in this logic- going from the fact that SAT is solvable by a LLM to the conclusion that the SAT is useless. Questions about addition can be solved using calculators, does that mean that teaching addition is useless? Also there is an underlying assumption that memorization is a useless skill, which I'd argue against, as in today's world most innovation is done by identifying patterns between discipline…
not so much memorization as decomposition of numerous examples into logic based knowledge. As a simple example, knowing volumes of words is not as useful as knowing etymology since etymology brings you a deeper structure.
Re: ChatGPT scores 80% on SAT reading/writing with collective chain of thought
#89Earlier quoted context omitted.
There is a leap in this logic- going from the fact that SAT is solvable by a LLM to the conclusion that the SAT is useless. Questions about addition can be solved using calculators, does that mean that teaching addition is useless? Also there is an underlying assumption that memorization is a useless skill, which I'd argue against, as in today's world most innovation is done by identifying patterns between discipline…
Well it shows that you can do well on SAT with basic word association without understanding and explaining the ideas of the passage. The "Joe Bloggs" theory has known this for ages.
Re: ChatGPT scores 80% on SAT reading/writing with collective chain of thought
#90Earlier quoted context omitted.
A human doing the same SATs will test roughly the same way every time. ChatGPT has achieved 80% in this user's run, will refuse to play at all the next run, will get 45% the following run etc. Stochastic randomisation makes the chatbot useful and provides these moments of impressive capability, it also means unlike humans that capability is in no way consistent.
ChatGPT isn't a human, it's a mix of data sampled from many humans. Random humans are often bad at SAT.
It's trained on the human data that was chosen as more impressive, and trained to pull the most impressive parts out of that data.
Fundamentally, while GPT may be "learned", it is not "logical". Instead of learning potential logical conclusions, it has learned potential semantic responses. It doesn't match the meaning of words to other meanings: it matchesa set of expressions to other expressions.
The impressive part is that the logical conclusions already present in the training data get presented by GPT in a semantically coherent way.