Live data from Hacker News

30% drop in O1-preview accuracy when Putnam problems are slightly variated

openreview.net

121–130 of 558 posts

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#121
post #92

The paper includes several examples of their modified questions. There has been a substantial jump from o1-preview to o1, so I gave several samples to o1 and o1-pro ( not o1-preview), and current o1s gave the correct answer to those modified problems. SOTA changes fast.

LLM boosters are so tiresome. You hardly did a rigorous evaluation, the set has been public since October and could have easily been added to the training data.

I'm not skilled enough in math to do a rigorous evaluation, so it was a quick check.

Terence Tao is skilled enough, and he describes O1's math ability is "...roughly on par with a mediocre, but not completely incompetent graduate student" (good discussion at https://news.ycombinator.com/item?id=41540902), and the next iteration O3 just got 25% on his brand new Frontier Math test.

Seeing LLMs as useless is banal, but downplaying their rate of improvement is self-sabotage.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#122
post #75

Link title says "slightly", but the PDF says two different kinds of variations: variable names (slight) and problem constants (significant), and the 30% drop is on the combination of a 26 variable and also 26 variable + constant questions. It's good to have a better test (though I bet this one will also be quickly saturated like all the others), but the title here doesn't seem justified by the page title there or the…

I would definitely classify both of those as slight changes. In fact I'd rename those as slight => trivial and significant => slight.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#123

So many negative comments as if o3 didn’t get 25% on frontiermath - which is absolutely nuts. Sure, LLMs will perform better if the answer to a problem is directly in their training set. But that doesn’t mean they perform bad when the answer isn’t in their training set.

EpochAI have to send the questions (but not the answer key) to OpenAI in order to score the models.

An overnight 2% -> 25% jump on this benchmark is a bit curious.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#124
post #110

Earlier quoted context omitted.

And what prompt would you give that does have novel input.

If I was me, I would start by giving a collection of LLMs the patent, ask half "why is this patent novel" and half "why is this patent not novel" and see what happens. I use this method of "debugging" my thinking (not code), might be a starting point here? Not sure.

Every patent application contains a section of claims. You can just ask the LLM to come up with ways to satisfy those claims.

But I'm sure there are lots of ways to go about it.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#125
post #84

Earlier quoted context omitted.

There is a reason why they won't do it. They are selling a narrative. There is a lot of money to be made here with this narrative and proving that artificial intelligence is NOT intelligent won't help sell that narrative.

The goal is to make it intelligent, by which OpenAI in particular explicitly mean "economically useful", not simply to be shiny. Passing tests is well known to be much easier than having deep understanding, even in humans. They openly ask for tests like this, not that they could possibly prevent them if they wanted to. There's scammers trying what you say of course, and I'm sure we've all seen some management initiat…

> The goal is to make it intelligent, by which OpenAI in particular explicitly mean "economically useful", not simply to be shiny

I never understood why this definition isn't a huge red flag for most people. The idea of boiling what intelligence is down to economic value is terrible, and inaccurate, in my opinion.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#126
The researcher's answer to their variant of "Year: 2016 ID: A1" in the appendix is wrong.

The solution (sum of 1,2,5,6,9,10,13,14, ...) has an alternating pattern, so has to be two piecewise interleaved polynomials, which cannot be expressed as a single polyomial.

Their answer works for k=1,2, but not k=3.

https://openreview.net/pdf?id=YXnwlZe0yf

This does not give me confidence in the results of their paper.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#127

Earlier quoted context omitted.

Because that is the whole conceit of how frontiermath is constructed

Didn't they run a bunch of models on the problem set? I doubt they are hosting all those models on their own infrastructure.

1. OpenAI has confirmed it’s not in their train (unlike putnam where they have never made any such claims)

2. They don't train on API calls

3. It is funny to me that HN finds it easier to believe theories about stealing data from APIs rather than an improvement in capabilities. It would be nice if symmetric scrutiny were applied to optimistic and pessimistic claims about LLMs, but I certainly don’t feel that is the case here.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#128
post #49

One experiment I would love to see, although not really feasible in practice, is to train a model on all digitized data from before the year 1905 (journals, letters, books, broadcasts, lectures, the works), and then ask it for a formula for mass-energy equivalence. A certain answer would definitely settle the debate on whether pattern recognition is a form of intelligence ;)

The best human performance on that task required many many hours of private work given that input.

How much would ChatGPT charge for that much reasoning? Isn't cost quadratic in sort term working memory?

It would be more interesting to prompt it with X% of a new paper's logical argument, and see if it can predict the rest.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#129
post #3

I hope someone reruns this on o1 and eventually o3. If o1-preview was the start like gpt1, then we should expect generalization to increase quickly.

I don't think llm generalise much, that's why they're not creative and can't solve novel problems. It's pattern matching with a huge amount of data. Study on the topic: https://arxiv.org/html/2406.15992v1 This would explain o1 poor performance with problems with variations. o3 seems to be expensive brute forcing in latent space followed by verification which should yield better results - but I don't think we can call…

From firsthand experience, this simply cannot be true. I can give them totally novel and unique physics problems I just made up- that requires tracking the movement of objects through a series of events, and it answers most correctly. Moreover, they find analogies between disparate concepts and fields of study and make useful suggestions based on them- which is arguably the same process as human creativity.

I think ultimately the disconnect is people theorizing about what it can or cannot do with an incorrect mental model of what it is, and then assuming it cannot do things that it can in fact do. The irony of discussions on LLMs is they more showcase the limits of humans ability to reason about novel situations.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#130

So many negative comments as if o3 didn’t get 25% on frontiermath - which is absolutely nuts. Sure, LLMs will perform better if the answer to a problem is directly in their training set. But that doesn’t mean they perform bad when the answer isn’t in their training set.

EpochAI have to send the questions (but not the answer key) to OpenAI in order to score the models. An overnight 2% -> 25% jump on this benchmark is a bit curious.

1. OpenAI said they did not train on these problems & they don’t train on API calls in general, that is a legal policy.

2. It was a new major model release from work over the course of months - struggle to see that as an ‘overnight’ jump in any real meaning.

3. Why is it easier to believe large scale corporate fraud than that the stated capabilities on a held out test set are real? Reads like cope, if I’m being frank.

Post reply on HN