Live data from Hacker News

30% drop in O1-preview accuracy when Putnam problems are slightly variated

openreview.net

1–10 of 558 posts

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#4
post #2

Is it just an open secret that the models are currently just being hardcoded for random benchmarks? Seems weird that people would be asking Putnam problems to a chatbot :/

Not hardcoded, I think it's just likely that those problems exist in its training data in some form

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#5
post #2

Is it just an open secret that the models are currently just being hardcoded for random benchmarks? Seems weird that people would be asking Putnam problems to a chatbot :/

Not hardcoded, I think it's just likely that those problems exist in its training data in some form

Isn't that just the LLM equivalent of hardcoding though?

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#6
post #2

Is it just an open secret that the models are currently just being hardcoded for random benchmarks? Seems weird that people would be asking Putnam problems to a chatbot :/

Not hardcoded, I think it's just likely that those problems exist in its training data in some form

Seems a bit picky. If the bot has seen the exact problem before it's not really doing anything more than recall to solve it.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#8
post #2

Is it just an open secret that the models are currently just being hardcoded for random benchmarks? Seems weird that people would be asking Putnam problems to a chatbot :/

Not hardcoded, I think it's just likely that those problems exist in its training data in some form

If I temember well this is call overfitting [1].

[1] https://en.wikipedia.org/wiki/Overfitting

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#10
post #2

Is it just an open secret that the models are currently just being hardcoded for random benchmarks? Seems weird that people would be asking Putnam problems to a chatbot :/

Not hardcoded, I think it's just likely that those problems exist in its training data in some form

I've always assumed they removed it, because it's such a basic and fundamental part of ML training that you separate your test and train data. And yet I never see any papers even mention if/how they do this. And I wonder if they do, how do they guarantee with high reliability that their massive terabytes of data don't contain the answer.
Post reply on HN