Live data from Hacker News

30% drop in O1-preview accuracy when Putnam problems are slightly variated

openreview.net

161–170 of 558 posts

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#161

Earlier quoted context omitted.

1. OpenAI has confirmed it’s not in their train (unlike putnam where they have never made any such claims) 2. They don't train on API calls 3. It is funny to me that HN finds it easier to believe theories about stealing data from APIs rather than an improvement in capabilities. It would be nice if symmetric scrutiny were applied to optimistic and pessimistic claims about LLMs, but I certainly don’t feel that is the c…

The modern state of training is to try to use everything they can get their hands on. Even if there are privileged channels that are guaranteed not to be used as training data, mentioning the problems on ancillary channels (say emailing another colleague to discuss the problem) can still create a risk of leakage because nobody making the decision to include the data is aware that stuff that should be excluded is in t…

Okay, then what about elite level codeforces performance? Those problems weren’t even constructed until after the model was made.

The real problem with all of these theories is most of these benchmarks were constructed after their training dataset cutoff points.

A sudden performance improvement on a new model release is not suspicious. Any model release that is much better than a previous one is going to be a “sudden jump in performance.”

Also, OpenAI is not reading your emails - certainly not with a less than one month lead time.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#162
post #2

Is it just an open secret that the models are currently just being hardcoded for random benchmarks? Seems weird that people would be asking Putnam problems to a chatbot :/

> Is it just an open secret that the models are currently just being hardcoded for random benchmarks? Seems weird that people would be asking Putnam problems to a chatbot :/

It's because people do keep asking these models math problems and then, when they get them right, citing it as evidence that they can actually do mathematical reasoning.

Since it's hard to determine what the models know, it's hard to determine when they're just spitting out something they were specifically trained on.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#163
post #118

Earlier quoted context omitted.

My wife is an 18th century American history professor. LLMs have very very clearly not been trained on 18th century English, they cannot really read it well, and they don't understand much from that period outside of very textbook stuff, anything nuanced or niche is totally missing. I've tried for over a year now, regularly, to help her use LLMs in her research, but as she very amusingly often says "your computers ar…

my wish for new years is that every time people make a comment like this they would share an example task

https://s.h4x.club/bLuNed45 - it's more crazy to me that my wife CAN in fact read this stuff easily, vs the fact that an LLM can't.

(for anyone who doesn't feel like downloading the zip, here is a single image from the zip: https://s.h4x.club/nOu485qx)

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#164

Earlier quoted context omitted.

Easier to believe or not, thinking that it's not a reasonable possibility is also funny.

Do you also think they somehow stole the codeforces problems before they were even written or you are willing to believe the #175 global rank there?

I dont think codeforce claims to contain novel unpublished problems.

But i'm not saying it's what they did, just that it's a possibility that should be considered till/if it is debunked.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#165
post #140

Earlier quoted context omitted.

> ask it for a formula for mass-energy equivalence Way too easy. If you think that mass and energy might be equivalent, then dimensional analysis doesn’t give you too much choice in the formula. Really, the interesting thing about E=mc^2 isn’t the formula but the assertion that mass is a form of energy and all the surrounding observations about the universe. Also, the actual insight in 1905 was more about asking the…

but e=mc^2 is just an approximation e: nice, downvoted for knowing special relativity

Can you elaborate? How is E=mc^2 an approximation, in special relativity or otherwise? What is it an approximation of?

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#166

Earlier quoted context omitted.

It certainly feels like certain patterns are hardcoded special cases, particularly to do with math. "Solve (1503+5171)*(9494-4823)" reliably gets the correct answer from ChatGPT "Write a poem about the solution to (1503+5171)*(9494-4823)" hallucinates an incorrect answer though That suggests to me that they've papered over the models inability to do basic math, but it's a hack that doesn't generalize beyond the simpl…

https://chatgpt.com/share/67755e6f-bfc8-8010-9aa3-8bcbbd9b26...

[deleted]

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#167

Earlier quoted context omitted.

It certainly feels like certain patterns are hardcoded special cases, particularly to do with math. "Solve (1503+5171)*(9494-4823)" reliably gets the correct answer from ChatGPT "Write a poem about the solution to (1503+5171)*(9494-4823)" hallucinates an incorrect answer though That suggests to me that they've papered over the models inability to do basic math, but it's a hack that doesn't generalize beyond the simpl…

https://chatgpt.com/share/67755e6f-bfc8-8010-9aa3-8bcbbd9b26...

To be clear I was testing with 4o, good to know that o1 has a better grasp of basic arithmetic. Regardless my point was less to do with the models ability to do math and more to do with OpenAI seeming to cover up its lack of ability.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#168
post #165

Earlier quoted context omitted.

but e=mc^2 is just an approximation e: nice, downvoted for knowing special relativity

Can you elaborate? How is E=mc^2 an approximation, in special relativity or otherwise? What is it an approximation of?

E^2 = m^2 + p^2 where p is momentum and i’ve dropped unit adjustment factors like c

this allows light to have energy even if its massless

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#169

Earlier quoted context omitted.

Didn't they run a bunch of models on the problem set? I doubt they are hosting all those models on their own infrastructure.

1. OpenAI has confirmed it’s not in their train (unlike putnam where they have never made any such claims) 2. They don't train on API calls 3. It is funny to me that HN finds it easier to believe theories about stealing data from APIs rather than an improvement in capabilities. It would be nice if symmetric scrutiny were applied to optimistic and pessimistic claims about LLMs, but I certainly don’t feel that is the c…

> 1. OpenAI has confirmed it’s not in their train (unlike putnam where they have never made any such claims)

Companies claim lots of things when it's in their best financial interest to spread that message. Unfortunately history has shown that in public communications, financial interest almost always trumps truth (pick whichever $gate you are aware of for convenience, i'll go with Dieselgate for a specific example).

> It is funny to me that HN finds it easier to believe theories about stealing data from APIs rather than an improvement in capabilities. It would be nice if symmetric scrutiny were applied to optimistic and pessimistic claims about LLMs, but I certainly don’t feel that is the case here.

What I see is generic unsubstantiated claims of artificial intelligence on one side and specific, reproducible examples that dismantle that claim on the other. I wonder how your epistemology works that leads you to accept marketing claims without evidence

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#170
post #163

Earlier quoted context omitted.

my wish for new years is that every time people make a comment like this they would share an example task

https://s.h4x.club/bLuNed45 - it's more crazy to me that my wife CAN in fact read this stuff easily, vs the fact that an LLM can't. (for anyone who doesn't feel like downloading the zip, here is a single image from the zip: https://s.h4x.club/nOu485qx )

have you been trying to provide it as an image directly? if so, doesn’t surprise me at all.

really thanks for sharing!

Post reply on HN