Live data from Hacker News

30% drop in O1-preview accuracy when Putnam problems are slightly variated

openreview.net

11–20 of 558 posts

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#11

This result is the same as a recent test of the same method+hypothesis from a group at Apple, no? I don’t have that reference handy but I don’t think I’m making it up.

I think you are probably referring to the following paper: https://arxiv.org/abs/2410.05229

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#15
post #3

I hope someone reruns this on o1 and eventually o3. If o1-preview was the start like gpt1, then we should expect generalization to increase quickly.

I don't think llm generalise much, that's why they're not creative and can't solve novel problems. It's pattern matching with a huge amount of data.

Study on the topic: https://arxiv.org/html/2406.15992v1

This would explain o1 poor performance with problems with variations. o3 seems to be expensive brute forcing in latent space followed by verification which should yield better results - but I don't think we can call it generalisation.

I think we need to go back to the drawing board.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#16

Performance of these LLMs on real life tasks feels very much like students last-minute cramming for Asian style exams. The ability to perfectly regurgitate, while no concept of meaning.

Basically yet another proof that we have managed to perfectly recreate human stupidity :-)

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#17
Or it's time to step back and call it what it is - very good pattern recognition.

I mean, that's cool... we can get a lot of work done with pattern recognition. Most of the human race never really moves above that level of thinking in the workforce or navigating their daily life, especially if they default to various societally prescribed patterns of getting stuff done (eg. go to college or the military , find a job , go to to find love....)

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#18

Performance of these LLMs on real life tasks feels very much like students last-minute cramming for Asian style exams. The ability to perfectly regurgitate, while no concept of meaning.

Basically yet another proof that we have managed to perfectly recreate human stupidity :-)

Good students are immune to variations that are discussed in the paper. But most academic tests may not differentiate between them and the crammers.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#19
post #3

I hope someone reruns this on o1 and eventually o3. If o1-preview was the start like gpt1, then we should expect generalization to increase quickly.

You're assuming that openAI isn't just gonna add the new questions to the training data.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#20

Or it's time to step back and call it what it is - very good pattern recognition. I mean, that's cool... we can get a lot of work done with pattern recognition. Most of the human race never really moves above that level of thinking in the workforce or navigating their daily life, especially if they default to various societally prescribed patterns of getting stuff done (eg. go to college or the military , find a job…

> Or it's time to step back and call it what it is - very good pattern recognition.

Or maybe it's time to stop wheeling out this tedious and disingenuous dismissal.

Saying it is just "pattern recognition" (or a "stochastic parrot") implies behavioural and performance characteristics that have very clearly been greatly exceeded.

Post reply on HN