Live data from Hacker News

ChatGPT scores 80% on SAT reading/writing with collective chain of thought

old.reddit.com

111–120 of 126 posts

Re: ChatGPT scores 80% on SAT reading/writing with collective chain of thought

#111
post #94

Earlier quoted context omitted.

Peak denialism. I and thousands of others use GPT's general creative intelligence, problem-solving skills and ability to manipulate abstract thoughts daily. Logically thinking the only way that ChatGPT could do well on a test of reasoning without engaging in reasoning is if the test is useless as a test of ability to reason. Since humans use this test to select people with good reasoning skills, why do you think it i…

> Peak denialism. That's a bit condescending. ChatGPT isn't remotely close to "general creative intelligence, problem-solving skills and ability to manipulate abstract thoughts daily". It only pretends, in a pretty impressive way. A bit like you can train your dog to play dead when you point your finger at them and say "bang", but you should not conclude from it that your dog understands the concept of death. > Since…

If you're not getting "general creative intelligence, problem-solving skills and ability to manipulate abstract thoughts daily" you're just not using ChatGPT right. I don't know what to tell you. It's like you are telling me dogs don't really fetch. My dog does. (Where fetch is an analogy for manipulating abstract thoughts.) There is no such thing as "pretending" to fetch, either the thing is bringing back the thought I threw or it's not. It brings it back no problem.

What are the kinds of things you are finding ChatGPT isn't able to understand?

Re: ChatGPT scores 80% on SAT reading/writing with collective chain of thought

#112
post #76

Earlier quoted context omitted.

Peak denialism. I and thousands of others use GPT's general creative intelligence, problem-solving skills and ability to manipulate abstract thoughts daily. Logically thinking the only way that ChatGPT could do well on a test of reasoning without engaging in reasoning is if the test is useless as a test of ability to reason. Since humans use this test to select people with good reasoning skills, why do you think it i…

Yeah, I use it too and I'm a big fan of ChatGPT and the other recent ML tools but I'm not under the illusion that this is intelligent machine. The test is not useless, the test's use case is to create a standardised datapoint to pick people to limited spots. That' what the test does, you need to attach false attributes to the test(like measures intelligence) to associate these attributes to the machine and this is a…

I've used ChatGPT a lot. It is an intelligent machine.

Re: ChatGPT scores 80% on SAT reading/writing with collective chain of thought

#113
post #108

Earlier quoted context omitted.

Peak denialism. I and thousands of others use GPT's general creative intelligence, problem-solving skills and ability to manipulate abstract thoughts daily. Logically thinking the only way that ChatGPT could do well on a test of reasoning without engaging in reasoning is if the test is useless as a test of ability to reason. Since humans use this test to select people with good reasoning skills, why do you think it i…

> test of reasoning My experience is that ChatGPT is not particularly good at reasoning, so I assume that either the test isn’t a particularly good test of reasoning or the result was a fluke. It’s a text model, not a reasoning model, so I wouldn’t expect reasoning ability in the first place.

You're right, it's not particularly good at reasoning.

Re: ChatGPT scores 80% on SAT reading/writing with collective chain of thought

#114
post #76

Earlier quoted context omitted.

Yeah, I use it too and I'm a big fan of ChatGPT and the other recent ML tools but I'm not under the illusion that this is intelligent machine. The test is not useless, the test's use case is to create a standardised datapoint to pick people to limited spots. That' what the test does, you need to attach false attributes to the test(like measures intelligence) to associate these attributes to the machine and this is a…

I've used ChatGPT a lot. It is an intelligent machine.

I guess if you define intelligence the “correct” way the calculators are intelligent. Anyway, what makes you think that you are in a unique position on having experience with chatGPT?

Re: ChatGPT scores 80% on SAT reading/writing with collective chain of thought

#115
post #107

GPT itself is in no way comparable to a person in terms of understanding . It makes no effort to determine the meaning of prompts or its training data. Instead, we are comparing the understanding of a person (who takes the SAT and scores something) against the logical soundness of a set of expressions written by other people (the training data), all weighted by GPT's ability to semantically transform those expression…

To me debating whether AI really understands anything is as irrelevant as asking whether submarines can swim . It performs tasks, and does them in its own way.

The full Edsger Dijkstra paragraph:

" The Fathers of the field had been pretty confusing: John von Neumann speculated about computers and the human brain in analogies sufficiently wild to be worthy of a medieval thinker and Alan M. Turing thought about criteria to settle the question of whether Machines Can Think, a question of which we now know that it is about as relevant as the question of whether Submarines Can Swim."

https://www.cs.utexas.edu/users/EWD/transcriptions/EWD08xx/E...

> It performs tasks, and does them in its own way.

The question is this: Are submarines 'artificial fish'?

(Dijkstra's typically cranky rant* was addressing something entirely different - note the title of the talk.)

* this humble HN user is simply channeling the iconoclastic spirit of the 'great man'.

Re: ChatGPT scores 80% on SAT reading/writing with collective chain of thought

#116
post #110
post #107

Earlier quoted context omitted.

To me debating whether AI really understands anything is as irrelevant as asking whether submarines can swim . It performs tasks, and does them in its own way.

Anyone know the Chinese Room experiment by Searle? GPT-3 was made for that :)

As discussed previously on HN: https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que...

Searle's thought experiment is due for an update. The entire force of the argument is simply to refute using Turing's Imitation game as a means of determining whether machines can think.

In Computing Machinery and Intelligence, Turing wrote:

"The new form of the problem can be described in terms of a game which we call the ‘imitation game’. It is played with three people, a man (A), a woman (B), and an interrogator (C) who may be of either sex. The interrogator stays in a room apart from the other two. The object of the game for the interrogator is to determine which of the other two is the man and which is the woman. He knows them by labels X and Y, and at the end of the game he says either ‘X is A and Y is B’ or ‘X is B and Y is A’. The interrogator is allowed to put questions to A and B thus: [question/answer follows ...]

"We now ask the question, ‘What will happen when a machine takes the part of A in this game?’ Will the interrogator decide wrongly as often when the game is played like this as he does when the game is played between a man and a woman? These questions replace our original, ‘Can machines think?’"

https://academic.oup.com/mind/article/LIX/236/433/986238

To misquote Djikstra, the Turing test is as relevant to question of machine intelligence as an underwater mobility test is to the question of whether a submarine is an artificial fish.

Re: ChatGPT scores 80% on SAT reading/writing with collective chain of thought

#117

Earlier quoted context omitted.

-- prompt -- Q: John has 5 bottles of water. He uses one every day. How many days before he runs out of water? A: John started with 5 bottles. He uses them at the rate of 1 bottle/day. 5 bottles divided by 1 bottle/day gives 5 days. The answer is 5 days. Q: Jill has 2 blankets. She uses 1 every night to keep warm. How many nights before she will sleep in the cold? -- response -- A: Jill started with 2 blankets. She u…

More, this time we will mix perishable and non-perishables and use CoK prompting again. -- prompt -- Q: Joe is a carpenter. He has 2 hammers and 5 bottles of water. In order to work, joe uses a hammer and one bottle every day. How many days can he work? A: Joe started with 2 hammers and 5 bottles. After 5 days he will exhaust his water bottles. The answer is 5 days. Q: Jill is an insomniac. She has 5 sleeping pills a…

-- prompt --

John posts 1 chatGPT challenge that is too ambiguous to be a valid test. James posts 14 well formed, unambiguous challenges that clearly highlight chatGPT's shortcomings. How many challenges remain too ambiguous to be valid?

-- response --

15 chairs.

Re: ChatGPT scores 80% on SAT reading/writing with collective chain of thought

#118
post #114

Earlier quoted context omitted.

I've used ChatGPT a lot. It is an intelligent machine.

I guess if you define intelligence the “correct” way the calculators are intelligent. Anyway, what makes you think that you are in a unique position on having experience with chatGPT?

My definition isn't so wide, most calculators don't manipulate abstract concepts in a way that makes me think they are an intelligent machine.

The reason I prefaced my opinion with "I've used ChatGPT a lot" is to show that this is my personal impression based on my real everyday experience - I'm not guessing or deducing. I personally interact with ChatGPT in a way that satisfies me as to the fact that it is an intelligent machine capable of very usefully handling abstract notions.

Re: ChatGPT scores 80% on SAT reading/writing with collective chain of thought

#119

I think this is evidence against the usefulness of standardized tests rather than support for ChatGPT

There is a leap in this logic- going from the fact that SAT is solvable by a LLM to the conclusion that the SAT is useless. Questions about addition can be solved using calculators, does that mean that teaching addition is useless? Also there is an underlying assumption that memorization is a useless skill, which I'd argue against, as in today's world most innovation is done by identifying patterns between discipline…

You're correct that memorization and pattern recognition are useful, but they're not supposed to be the majority of the test!

A good test should also include reasoning skills and comprehension.

Having reasoning skills means not all necessary information can be derived from the test itself, the necessary information isn't always well defined, yet a correct answer is produced.

Demonstrating comprehension requires the test taker to write out a model of what they think they know in their own words. Undoubtedly the answer will be at least slightly wrong which is why subjective expertise is required to evaluate the level of comprehension.

To me it seems the quality of the tests has deteriorated due to cost cutting, not because AI has become so much more sophisticated.

Subjectivity is usually dismissed these days as a roll of the dice, but it's vital to fitting the real world full of humans living in it. Humans must be the gatekeepers, not to fight off AI, but by the very definition of what it means to be convincing.

Answers to real world questions solve many more constraints than just the concrete subject matter. The whole point to updating a test on a well understood topic should not merely be to prevent "replay attack" cheating, but to conform to the new subjectivity on the topic.

Re: ChatGPT scores 80% on SAT reading/writing with collective chain of thought

#120
post #27

Earlier quoted context omitted.

The AP CSA pass surprises me (I've never taken it myself). I had this thought because I tried to use chatgpt for work programming stuff and it got stuff wrong very frequently. So frequently, and to a great enough degree, that I had to keep re-circumscribing what kind of problems were worth my time to throw at it, and now it's a tiny cross-section.

Thats what surprises me as well. If I give chatGPT anything that requires actual thought or reasoning put into it, it just produces gibberish. It's great for getting a blueprint to start from. I recently had to write a gtk+rust client, and I asked chatGPT to give me some starter code which largely worked. All the chatGPT hype feels the same to me as the copilot hype from a while back. The people who got the most exci…

That does seem like a good application for it. If one thinks something is in the corpus many times (various GTK idioms), chatgpt is pretty good at a type of syntactic composition...e.g., "can you show it to me in rust"
Post reply on HN