Live data from Hacker News

30% drop in O1-preview accuracy when Putnam problems are slightly variated

openreview.net

321–330 of 558 posts

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#321

Earlier quoted context omitted.

FYI: If you do that without a subscrpition, you currently (most likely) get a response generated through 4o-mini — which is not any of their reasoning models (o1, o1-mini or previously o1-preview) of the branch discussed in the linked paper. Notably, it's not even necessarily 4o, their premiere "non-reasoning"-model, but likely the cheaper variant: With a free account the model it claims to be using is "4o auto", whi…

There doesn't seem to be a way to choose a model up-front with a free account, but after you make a query you can click on the "regenerate" button and select whether to try again with "auto", 4o, or 4o-mini. At least until you use 4o too many times and get rate limited.

you can select the model in the header bar when you start a chat: the name of the currently selected model can be clicked to reveal a dropdown

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#322
post #197

I remember when this stuff was all coming out and people were finally excited about ChatGPT getting the problem with "which is heavier, a 10 pound bag of feathers or a 10 pound bag of bricks?" problem correct. But of course it got it correct. It was in the training set. Vary the problem slightly by just changing the nouns, or changing the numbers so that one in fact was heavier than the other, and performance went al…

ChatGPT Plus user here. The following are all fresh sessions and first answers, no fishing. GPT 4: The 10.01-pound bag of fluffy cotton is heavier than the 9.99-pound bag of steel ingots. The type of material doesn’t affect the weight comparison; it’s purely a matter of which bag weighs more on the scale. GPT 4o: The 10.01-pound bag of fluffy cotton is heavier. Weight is independent of the material, so the bag of cot…

they've likely read this thread and adjusted their pre-filter to give the correct answer

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#323

Earlier quoted context omitted.

> give me a query and i’ll ask it Which is heavier: an 11kg bag of lint or a 20lb bag of gold?

yeah it gets it https://chatgpt.com/share/67757720-3c7c-8010-a3e9-ce66fb9f17... e: cool, this gets downvoted

It got it right, but an interesting result that it rambled on about monetary value for... no reason.

> While the lint bag is heavier in terms of weight, it's worth mentioning that gold is significantly more valuable per pound compared to lint. This means that even though the lint bag weighs more, the gold bag holds much greater monetary value.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#324
post #152

Earlier quoted context omitted.

> The goal is to make it intelligent, by which OpenAI in particular explicitly mean "economically useful", not simply to be shiny I never understood why this definition isn't a huge red flag for most people. The idea of boiling what intelligence is down to economic value is terrible, and inaccurate, in my opinion.

Everyone has a very different idea of what the word "intelligence" means; this definition has got the advantage that, unlike when various different AI became superhuman at arithmetic, symbolic logic, chess, jeopardy, go, poker, number of languages it could communicate in fluently, etc., it's tied to tasks people will continuously pay literally tens of trillions of dollars each year for because they want those tasks d…

Maybe by the time it’s doing a trillion dollars a year of useful work (less than 10 years out) people will call it intelligent… but still probably not.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#325
post #204

Earlier quoted context omitted.

https://chatgpt.com/share/67756897-8974-8010-a0e0-c9e3b3e91f... so far o1-mini has bodied every task people are saying LLMs can’t do in this thread

That appears to be the same model I used. This is why I emphasized I didn't "go shopping" for a result. That was the first result I got. I'm not at all surprised that it will nondeterministically get it correct sometimes. But if it doesn't get it correct every time, it doesn't "know". (In fact "going shopping" for errors would still even be fair. It should be correct all the time if it "knows". But it would be differ…

Could you share the exact chat you used for when it failed? There is a share chat button on openai.

It's very difficult to be an AI bull when the goalposts are moving so quickly that ai answering core correctly across multiple models is brushed off as 'nondeterministically getting it correct sometimes'

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#326

Earlier quoted context omitted.

There doesn't seem to be a way to choose a model up-front with a free account, but after you make a query you can click on the "regenerate" button and select whether to try again with "auto", 4o, or 4o-mini. At least until you use 4o too many times and get rate limited.

you can select the model in the header bar when you start a chat: the name of the currently selected model can be clicked to reveal a dropdown

That option isn't there for me, maybe it's an A/B test thing.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#327

Yes so when you change the sequence of tokens they've electronically memorized, they get a bit worse at predicting the next token?

When you put it that way it’s a trivial result. However the consequences for using AI to replace humans on tasks is significant.

The only people super pumping the idea of mass replacement of human labor are financially invested in that outcome.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#328
post #197

I remember when this stuff was all coming out and people were finally excited about ChatGPT getting the problem with "which is heavier, a 10 pound bag of feathers or a 10 pound bag of bricks?" problem correct. But of course it got it correct. It was in the training set. Vary the problem slightly by just changing the nouns, or changing the numbers so that one in fact was heavier than the other, and performance went al…

ChatGPT Plus user here. The following are all fresh sessions and first answers, no fishing. GPT 4: The 10.01-pound bag of fluffy cotton is heavier than the 9.99-pound bag of steel ingots. The type of material doesn’t affect the weight comparison; it’s purely a matter of which bag weighs more on the scale. GPT 4o: The 10.01-pound bag of fluffy cotton is heavier. Weight is independent of the material, so the bag of cot…

So do what the commenter suggests and make irrelevant permutations to the input to find when it fails. ie., engage in hypothesis testing rather than confirmation bias.

If a system has the capability to solve problems of {parts1...parts_n}, then it only has that capability if irrelevant permutations {parts1..parts2'...parts_n} make no difference.

Its very obvious that such permutations can destory such apparent capabilities.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#329

Earlier quoted context omitted.

No, it only says "ChatGPT Plus (Upgrade)" or "ChatGPT". Maybe it's different if you have a paid account?

if i go to chatgpt.com on my phone not logged on at all it tells me very prominently at the top that i am using 4o mini

Logged in, non paid account, on a desktop, for me, it's exactly as the person you're replying to has stated.

If I log out, it shows 4o mini, and when I try to change it, it asks me to login or sign in rather than giving me any options.

When I use enough chatgpt when logged in it gives me some nebulous "you've used all your xyz tokens for the day". But other than that there is no real signal to me that I'm getting a degraded experience.

It's really just confusing as hell.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#330

Earlier quoted context omitted.

ChatGPT Plus user here. The following are all fresh sessions and first answers, no fishing. GPT 4: The 10.01-pound bag of fluffy cotton is heavier than the 9.99-pound bag of steel ingots. The type of material doesn’t affect the weight comparison; it’s purely a matter of which bag weighs more on the scale. GPT 4o: The 10.01-pound bag of fluffy cotton is heavier. Weight is independent of the material, so the bag of cot…

So do what the commenter suggests and make irrelevant permutations to the input to find when it fails. ie., engage in hypothesis testing rather than confirmation bias. If a system has the capability to solve problems of {parts1...parts_n}, then it only has that capability if irrelevant permutations {parts1..parts2'...parts_n} make no difference. Its very obvious that such permutations can destory such apparent capabi…

> ...engage in hypothesis testing rather than confirmation bias

Please leave the premises, sir. We don't take kindly to luddites here.

Post reply on HN