Live data from Hacker News

30% drop in O1-preview accuracy when Putnam problems are slightly variated

openreview.net

91–100 of 558 posts

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#91
It drops from 50 to 33,96. Still the best, o1 on variable is around 2 times better than Claude on original test.

The rest of the llm are far away, single digit.

It makes me wonder if o1 is finally getting intelligent? LLM are not supposed to understand these problems when you change variable and values, they have to rely on preexisting data of absolutely identical solved problem to give a correct answer.

I didn't follow LLM development but I heard one times that chatgpt is now composed of multiple LLM and maybe they put multiple artificial intelligence with purpose of problems solvings or trigonometry for instance.

That would explain the reason it's so much better.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#92
The paper includes several examples of their modified questions. There has been a substantial jump from o1-preview to o1, so I gave several samples to o1 and o1-pro (not o1-preview), and current o1s gave the correct answer to those modified problems. SOTA changes fast.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#93
post #49

One experiment I would love to see, although not really feasible in practice, is to train a model on all digitized data from before the year 1905 (journals, letters, books, broadcasts, lectures, the works), and then ask it for a formula for mass-energy equivalence. A certain answer would definitely settle the debate on whether pattern recognition is a form of intelligence ;)

This reminds me of a similar idea I recently heard in podcast with Adam Brown. I'm unsure whether it is his original notion. The idea being, that if we can create AI that can derive special relativity (1905) from pre-Einstein books and papers then we have reached the next game-changing milestone in the advancement of artificial reasoning.

Great podcast, especially the part about hitchhiking :)

https://www.youtube.com/watch?v=XhB3qH_TFds

Or RSS

https://api.substack.com/feed/podcast/69345.rss

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#94
post #64

Earlier quoted context omitted.

I think what it shows that it has minimal "understanding" of the problem - otherwise such small variations wouldn't pose a challenge. Training it to handle these specific small variations doesn't change that. It's good in automation, not understanding.

If it were a complete failure on variations I would be inclined to agree. Instead it was a 30% drop in performance. I would characterise that as limited understanding.

My guess is that what’s understood isn’t various parts of solving the problem but various aspects of the expected response.

I see this more akin to a human faking their way through a conversation.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#95
post #48

Earlier quoted context omitted.

I've always assumed they removed it, because it's such a basic and fundamental part of ML training that you separate your test and train data. And yet I never see any papers even mention if/how they do this. And I wonder if they do, how do they guarantee with high reliability that their massive terabytes of data don't contain the answer.

I don't see any reason to assume they removed it unless they're very explicit about it. Model publishers have an extremely strong vested interest in beating benchmarks and I expect them to teach to the test if they can get away with it.

As usual, once a metric becomes a target, it stops being useful.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#96
post #43
post #22

Earlier quoted context omitted.

Citation needed. Please be more specific, or else this is just a tedious and disingenuous advocacy.

Gpt4 can add very large integers. It is evident that it is not recalling the sum because all combinations of integer addition were likely not in the training data, Storing the answer to the sum of all integers up to the size that GPT4 can manage would take more parameters than the model has. That addition is a small capability but you only need a single counterexample to disprove a theory.

> That addition is a small capability but you only need a single counterexample to disprove a theory

No, that's not how this works :)

You can hardcode an exception to pattern recognition for specific cases - it doesn't cease to be a pattern recognizer with exceptions being sprinkled in.

The 'theory' here is that a pattern recognizer can lead to AGI. That is the theory. Someone saying 'show me proof or else I say a pattern recognizer is just a pattern recognizer' is not a theory and thus cannot be disproven, or proven.

This is also known as Russell's teapot. https://en.wikipedia.org/wiki/Russell%27s_teapot

If someone claims there's a teapot out in space - the burden of proof is on the person making the claim, not on the person saying it is bullshit.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#97
post #92

The paper includes several examples of their modified questions. There has been a substantial jump from o1-preview to o1, so I gave several samples to o1 and o1-pro ( not o1-preview), and current o1s gave the correct answer to those modified problems. SOTA changes fast.

LLM boosters are so tiresome. You hardly did a rigorous evaluation, the set has been public since October and could have easily been added to the training data.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#98
post #49

One experiment I would love to see, although not really feasible in practice, is to train a model on all digitized data from before the year 1905 (journals, letters, books, broadcasts, lectures, the works), and then ask it for a formula for mass-energy equivalence. A certain answer would definitely settle the debate on whether pattern recognition is a form of intelligence ;)

But is there even enough pre-1905 data to create models that say hello world reliably?

The terabytes of training data required for decent LLMs does not exist. I’d guess there may only be gigabytes worth.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#99
post #49

One experiment I would love to see, although not really feasible in practice, is to train a model on all digitized data from before the year 1905 (journals, letters, books, broadcasts, lectures, the works), and then ask it for a formula for mass-energy equivalence. A certain answer would definitely settle the debate on whether pattern recognition is a form of intelligence ;)

There is a reason why they won't do it. They are selling a narrative. There is a lot of money to be made here with this narrative and proving that artificial intelligence is NOT intelligent won't help sell that narrative.

They don't have to do it themselves. The super-GPU cluster used to train GPT-6 will eventually shrink down to a garage size and eventually some YouTuber will.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#100
post #61
post #49

One experiment I would love to see, although not really feasible in practice, is to train a model on all digitized data from before the year 1905 (journals, letters, books, broadcasts, lectures, the works), and then ask it for a formula for mass-energy equivalence. A certain answer would definitely settle the debate on whether pattern recognition is a form of intelligence ;)

This is how patent disputes should be decided. If an LLM can figure it out, then it is not novel.

And what prompt would you give that does have novel input.
Post reply on HN