Live data from Hacker News

30% drop in O1-preview accuracy when Putnam problems are slightly variated

openreview.net

341–350 of 558 posts

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#341
This is very interesting, but a couple of things to note; 1. o1 still achieves > 40% on the varied Putnam problems, which is still a feat most math students would not achieve. 2. o3 solved 25% of the Epoch AI dataset. - There was an interesting post which calls into question how difficult some of those problems actually are, but it still seems very impressive.

I think a fair conclusion here is reasoning models are still really good at solving very difficult math and competitive programming problems, but just better at ones they have seen before.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#342

Earlier quoted context omitted.

give me a query and i’ll ask it, but also i don’t want to burn through all of my o1mini allocation and have to use the pay-as-you-go API.

>> so far o1-mini has bodied every task people are saying LLMs can’t do in this thread > give me a query and i’ll ask it Here's a query similar to one that I gave to Google Gemini (version unknown), which failed miserably: ---query--- Steeleye Span's version of the old broadsheet ballad "The Victory" begins the final verse with these lines: Here's success unto the Victory / and crew of noble fame and glory to the cap…

i'd prefer an easily verifiable question rather than one where we can always go "no that's not what they really meant" but someone else with o1-mini quota can respond

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#343

Earlier quoted context omitted.

ChatGPT Plus user here. The following are all fresh sessions and first answers, no fishing. GPT 4: The 10.01-pound bag of fluffy cotton is heavier than the 9.99-pound bag of steel ingots. The type of material doesn’t affect the weight comparison; it’s purely a matter of which bag weighs more on the scale. GPT 4o: The 10.01-pound bag of fluffy cotton is heavier. Weight is independent of the material, so the bag of cot…

So do what the commenter suggests and make irrelevant permutations to the input to find when it fails. ie., engage in hypothesis testing rather than confirmation bias. If a system has the capability to solve problems of {parts1...parts_n}, then it only has that capability if irrelevant permutations {parts1..parts2'...parts_n} make no difference. Its very obvious that such permutations can destory such apparent capabi…

If GP's hypothesis was "it fails for small variations of the input, like this one", then testing that hypothesis with that exact variation on a couple models seems fair and scientific.

Testing it with more variations until one fails feels a bit like p-hacking. You'd need to engage in actual statistics to get reliable results from that, beyond "If I really try, I can make it fail". Which would be a completely different hypothesis than the one presented at the start

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#344
post #197

I remember when this stuff was all coming out and people were finally excited about ChatGPT getting the problem with "which is heavier, a 10 pound bag of feathers or a 10 pound bag of bricks?" problem correct. But of course it got it correct. It was in the training set. Vary the problem slightly by just changing the nouns, or changing the numbers so that one in fact was heavier than the other, and performance went al…

Are you sure you weren't fishing? I ran 5 sessions and never got the wrong answer. All using gpt 4o-mini, which is the default non logged in experience on chatgpt.com. 1. The 10.01-pound bag of fluffy cotton is heavier than the 9.99-pound bag of steel ingots. Despite the difference in material density, the key factor here is the weight itself, with 10.01 pounds being greater than 9.99 pounds, regardless of the substa…

Not OP, but I got 4o-mini confused on second attempt.

https://chatgpt.com/share/67759d1a-1430-800b-a0a9-2c5f2ac02a...

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#345
post #197

I remember when this stuff was all coming out and people were finally excited about ChatGPT getting the problem with "which is heavier, a 10 pound bag of feathers or a 10 pound bag of bricks?" problem correct. But of course it got it correct. It was in the training set. Vary the problem slightly by just changing the nouns, or changing the numbers so that one in fact was heavier than the other, and performance went al…

Add some extra information, and it gets confused. This is 4o. https://chatgpt.com/share/67759723-f008-800e-b0f3-9c81e656d6... One might argue that it's impossible to compress air using known engineering, but that would be a different kind of answer.

It seems more like ChatGPT was asked a rather bizarre question with far too little detail to make sense, and ChatGPT failed to notice or to ask for more information. Although it did get rather impressively confused about the pressure of the air.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#346

Earlier quoted context omitted.

> ...engage in hypothesis testing rather than confirmation bias Please leave the premises, sir. We don't take kindly to luddites here.

Tough crowd

Lots of other websites are more appropriate for meme jokes.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#347

Earlier quoted context omitted.

So do what the commenter suggests and make irrelevant permutations to the input to find when it fails. ie., engage in hypothesis testing rather than confirmation bias. If a system has the capability to solve problems of {parts1...parts_n}, then it only has that capability if irrelevant permutations {parts1..parts2'...parts_n} make no difference. Its very obvious that such permutations can destory such apparent capabi…

If GP's hypothesis was "it fails for small variations of the input, like this one", then testing that hypothesis with that exact variation on a couple models seems fair and scientific. Testing it with more variations until one fails feels a bit like p-hacking. You'd need to engage in actual statistics to get reliable results from that, beyond "If I really try, I can make it fail". Which would be a completely differen…

Except that if the model genuinely was reasoning about the problem, you could test it with every variation of materials and weights in the world and it would pass. Failing that problem at all in any way under any conditions is a failure of reasoning.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#348
post #204

Earlier quoted context omitted.

That appears to be the same model I used. This is why I emphasized I didn't "go shopping" for a result. That was the first result I got. I'm not at all surprised that it will nondeterministically get it correct sometimes. But if it doesn't get it correct every time, it doesn't "know". (In fact "going shopping" for errors would still even be fair. It should be correct all the time if it "knows". But it would be differ…

Could you share the exact chat you used for when it failed? There is a share chat button on openai. It's very difficult to be an AI bull when the goalposts are moving so quickly that ai answering core correctly across multiple models is brushed off as 'nondeterministically getting it correct sometimes'

Why? Did a grocery store self checkout ever fail to calculate sales tax? Do I need to run a study on that?

The people selling this could not make a car drive but now its AGI.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#349

Earlier quoted context omitted.

So do what the commenter suggests and make irrelevant permutations to the input to find when it fails. ie., engage in hypothesis testing rather than confirmation bias. If a system has the capability to solve problems of {parts1...parts_n}, then it only has that capability if irrelevant permutations {parts1..parts2'...parts_n} make no difference. Its very obvious that such permutations can destory such apparent capabi…

If GP's hypothesis was "it fails for small variations of the input, like this one", then testing that hypothesis with that exact variation on a couple models seems fair and scientific. Testing it with more variations until one fails feels a bit like p-hacking. You'd need to engage in actual statistics to get reliable results from that, beyond "If I really try, I can make it fail". Which would be a completely differen…

I feel like I'm almost 100% certain that the smart guys at OpenAI have added many more variations of the problem to their training set since OP did his failing test, so it doesn't surprise me at all to know that this exact one now passes.

In fact, in my use of o1 it's incredibly clear that it still has the same problems. It's incredibly common that the second I ask for someone even slightly outside the training set, it's more likely to "round" to some wrong solution in the training set, rather than use any sort of human-like reasoning to figure out the right answer (often the right answer isn't hard to get, just not found in a Google search).

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#350
post #332

Earlier quoted context omitted.

This happens literally every time. Someone always says "ChatGPT can't do this!", but then when someone actually runs the example, chatGPT gets it right. Now what the OP is going to do next is proceed to move goalposts and say like "but umm I just asked chatgpt this, so clearly they modified the code in realtime to get the answer right"

Prompt: In the Netherlands, in terms of drinks, is there a particular spirit that represents the country? > Yes, in the Netherlands, jenever (also known as genever) is the traditional spirit that represents the country. Jenever is a type of Dutch gin that has a distinctive flavor, often made from malt wine and flavored with juniper berries. It has a long history in the Netherlands, dating back to the 16th century, an…

'Berenberg is made by adding herbs to jenever'

From your comment it would seem that you are disputing jenever's popularity by saying jenever is more popular...

Perhaps it was a good faith mistake? If so, that would imply that the AI knows more about jenever than you?

Post reply on HN