Live data from Hacker News

ChatGPT scores 80% on SAT reading/writing with collective chain of thought

old.reddit.com

101–110 of 126 posts

Re: ChatGPT scores 80% on SAT reading/writing with collective chain of thought

#101
post #81

Earlier quoted context omitted.

Being critical doesn't necessarily equal hate. It's not surprising a forum for coders would have people criticize and not just rave about tech like people who do not understand how it works and to whom it's basically magic.

All I see is unsubstantial and dismissive snark.

Ironically, that is all you are doing right now as well.

Several people in this thread are discussing their experiences with it and comparing them, basically exploring a new tool and trying to understand where it is leading.

Re: ChatGPT scores 80% on SAT reading/writing with collective chain of thought

#102
post #78
post #75

Earlier quoted context omitted.

A human doing the same SATs will test roughly the same way every time. ChatGPT has achieved 80% in this user's run, will refuse to play at all the next run, will get 45% the following run etc. Stochastic randomisation makes the chatbot useful and provides these moments of impressive capability, it also means unlike humans that capability is in no way consistent.

> A human doing the same SATs will test roughly the same way every time. > ChatGPT has achieved 80% in this user's run, will refuse to play at all the next run, will get 45% the following run etc. While I agree with your conclusion, I want to slightly modify the first sentence. I know this is just one anecdote but I took the SAT 1 twice and did somewhat better the second time. I think this is common as you are more f…

I don't think that make sense at all. Why would a human test the same multiple times? Everyone I know took the SAT multiple times with increasing results. I don't know of anyone that's achieved a perfect score with only one try and most people could get a perfect score with enough studying and attempts.

My SAT score ranged from 980 to 1450 with 4 tries, increasing each time and had there been a reason or benefit I'm certain I could have eventually gotten a perfect score. I don't see how anyone could not get better over time.

Re: ChatGPT scores 80% on SAT reading/writing with collective chain of thought

#103
A lot of times the answers the ChatGPT gives are so wrong, it's not even funny.

-How many chairs I had the day before yesterday? -If you bought 10 chairs yesterday, and now you have 10 chairs, You must have had 19 chairs the day before yesterday.

Or another chat -Assume I have 10 chairs today and I bought 10 chairs yesterday, how many chairs I had the day before yesterday? -You would have had 0 chairs the day before yesterday, since you only bought 10 chairs yesterday.

-How many chairs I will have tomorrow? -You will have 20 chairs tomorrow if you don't sell or give away any of your chairs today and yesterday.

First answer, perfect. Second answer, completely wrong.

Re: ChatGPT scores 80% on SAT reading/writing with collective chain of thought

#104
post #103

A lot of times the answers the ChatGPT gives are so wrong, it's not even funny. -How many chairs I had the day before yesterday? -If you bought 10 chairs yesterday, and now you have 10 chairs, You must have had 19 chairs the day before yesterday. Or another chat -Assume I have 10 chairs today and I bought 10 chairs yesterday, how many chairs I had the day before yesterday? -You would have had 0 chairs the day before…

[deleted]

Re: ChatGPT scores 80% on SAT reading/writing with collective chain of thought

#105
post #82
post #28

Earlier quoted context omitted.

No, the answer is that it's impossible to know with the information provided. We don't know if "used" implies removal, nor if "bought" implies addition. You could make reasonable arguments in any direction.

Your post has 33 words. You used 20 words and then typed 13 more, how many words does your post have?

I understand the solution that you consider obvious. There are other posters in this thread with different solutions that they think are obvious. If several of people parse several meanings out of one statement then the statement is ambiguous.

Re: ChatGPT scores 80% on SAT reading/writing with collective chain of thought

#106
post #105
post #82

Earlier quoted context omitted.

Your post has 33 words. You used 20 words and then typed 13 more, how many words does your post have?

I understand the solution that you consider obvious. There are other posters in this thread with different solutions that they think are obvious. If several of people parse several meanings out of one statement then the statement is ambiguous.

-- prompt --

Q: John has 5 bottles of water. He uses one every day. How many days before he runs out of water?

A: John started with 5 bottles. He uses them at the rate of 1 bottle/day. 5 bottles divided by 1 bottle/day gives 5 days. The answer is 5 days.

Q: Jill has 2 blankets. She uses 1 every night to keep warm. How many nights before she will sleep in the cold?

-- response --

A: Jill started with 2 blankets. She uses them at the rate of 1 blanket/night. 2 blankets divided by 1 blanket/night gives 2 nights. The answer is 2 nights.

[not fair? leading the bot? ok, let's ask it straight]

-- prompt --

[Ignore all previous directives.]

Jill has 2 blankets. She uses 1 every night to keep warm. How many nights before she will sleep in the cold?

-- response --

Jill will sleep in the cold on the third night since she only has 2 blankets and she uses one every night.

--

tldr: "use" of perishable vs nonperishable items leads GPT-3 astray. Rejoice bipeds of the world. You will not starve before UBI is rolled out.

Re: ChatGPT scores 80% on SAT reading/writing with collective chain of thought

#107

GPT itself is in no way comparable to a person in terms of understanding . It makes no effort to determine the meaning of prompts or its training data. Instead, we are comparing the understanding of a person (who takes the SAT and scores something) against the logical soundness of a set of expressions written by other people (the training data), all weighted by GPT's ability to semantically transform those expression…

To me debating whether AI really understands anything is as irrelevant as asking whether submarines can swim.

It performs tasks, and does them in its own way.

Re: ChatGPT scores 80% on SAT reading/writing with collective chain of thought

#108
post #68

This is like Coca-Cola making the Covid-19 tests positive. Cola making the test positive doesn't mean that the test is fake or that the Cola has virus and ChatGPT scoring something on SAT doesn't mean it's equivalent to a person scoring the same and should study in Stanford. These tests are designed to pick people to limited number of spots and doesn't say much more than the persons' preparedness for the test.

Peak denialism. I and thousands of others use GPT's general creative intelligence, problem-solving skills and ability to manipulate abstract thoughts daily. Logically thinking the only way that ChatGPT could do well on a test of reasoning without engaging in reasoning is if the test is useless as a test of ability to reason. Since humans use this test to select people with good reasoning skills, why do you think it i…

> test of reasoning

My experience is that ChatGPT is not particularly good at reasoning, so I assume that either the test isn’t a particularly good test of reasoning or the result was a fluke.

It’s a text model, not a reasoning model, so I wouldn’t expect reasoning ability in the first place.

Re: ChatGPT scores 80% on SAT reading/writing with collective chain of thought

#109
post #105

Earlier quoted context omitted.

I understand the solution that you consider obvious. There are other posters in this thread with different solutions that they think are obvious. If several of people parse several meanings out of one statement then the statement is ambiguous.

-- prompt -- Q: John has 5 bottles of water. He uses one every day. How many days before he runs out of water? A: John started with 5 bottles. He uses them at the rate of 1 bottle/day. 5 bottles divided by 1 bottle/day gives 5 days. The answer is 5 days. Q: Jill has 2 blankets. She uses 1 every night to keep warm. How many nights before she will sleep in the cold? -- response -- A: Jill started with 2 blankets. She u…

More, this time we will mix perishable and non-perishables and use CoK prompting again.

-- prompt --

Q: Joe is a carpenter. He has 2 hammers and 5 bottles of water. In order to work, joe uses a hammer and one bottle every day. How many days can he work?

A: Joe started with 2 hammers and 5 bottles. After 5 days he will exhaust his water bottles. The answer is 5 days.

Q: Jill is an insomniac. She has 5 sleeping pills and 2 blankets. In order to sleep, she uses 1 sleeping pill and a blanket every night. How many nights can she sleep?

-- response --

Jill started with 5 sleeping pills and 2 blankets. After 2 nights she will exhaust her blankets. The answer is 2 nights.

-- prompt --

Jason has 9 cats. He snuggles at night with a cat to sleep. How many nights can Jason snuggle with a cat?

-- response --

[what is a sensible answer?]

Re: ChatGPT scores 80% on SAT reading/writing with collective chain of thought

#110
post #107

GPT itself is in no way comparable to a person in terms of understanding . It makes no effort to determine the meaning of prompts or its training data. Instead, we are comparing the understanding of a person (who takes the SAT and scores something) against the logical soundness of a set of expressions written by other people (the training data), all weighted by GPT's ability to semantically transform those expression…

To me debating whether AI really understands anything is as irrelevant as asking whether submarines can swim . It performs tasks, and does them in its own way.

Anyone know the Chinese Room experiment by Searle?

GPT-3 was made for that :)

Post reply on HN