This result is the same as a recent test of the same method+hypothesis from a group at Apple, no? I don’t have that reference handy but I don’t think I’m making it up.
30% drop in O1-preview accuracy when Putnam problems are slightly variated
11–20 of 558 posts
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#12I would love to see how well Deepseek V3 do on this.
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#13The ability to perfectly regurgitate, while no concept of meaning.
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#14If you are shocked by this, you are the sucker in the room.
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#15I hope someone reruns this on o1 and eventually o3. If o1-preview was the start like gpt1, then we should expect generalization to increase quickly.
Study on the topic: https://arxiv.org/html/2406.15992v1
This would explain o1 poor performance with problems with variations. o3 seems to be expensive brute forcing in latent space followed by verification which should yield better results - but I don't think we can call it generalisation.
I think we need to go back to the drawing board.
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#16Performance of these LLMs on real life tasks feels very much like students last-minute cramming for Asian style exams. The ability to perfectly regurgitate, while no concept of meaning.
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#17I mean, that's cool... we can get a lot of work done with pattern recognition. Most of the human race never really moves above that level of thinking in the workforce or navigating their daily life, especially if they default to various societally prescribed patterns of getting stuff done (eg. go to college or the military , find a job , go to to find love....)
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#18Performance of these LLMs on real life tasks feels very much like students last-minute cramming for Asian style exams. The ability to perfectly regurgitate, while no concept of meaning.
Basically yet another proof that we have managed to perfectly recreate human stupidity :-)
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#19I hope someone reruns this on o1 and eventually o3. If o1-preview was the start like gpt1, then we should expect generalization to increase quickly.
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#20Or it's time to step back and call it what it is - very good pattern recognition. I mean, that's cool... we can get a lot of work done with pattern recognition. Most of the human race never really moves above that level of thinking in the workforce or navigating their daily life, especially if they default to various societally prescribed patterns of getting stuff done (eg. go to college or the military , find a job…
Or maybe it's time to stop wheeling out this tedious and disingenuous dismissal.
Saying it is just "pattern recognition" (or a "stochastic parrot") implies behavioural and performance characteristics that have very clearly been greatly exceeded.