Performance of these LLMs on real life tasks feels very much like students last-minute cramming for Asian style exams. The ability to perfectly regurgitate, while no concept of meaning.
30% drop in O1-preview accuracy when Putnam problems are slightly variated
101–110 of 558 posts
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#102Earlier quoted context omitted.
Good students are immune to variations that are discussed in the paper. But most academic tests may not differentiate between them and the crammers.
> Good students are immune to variations I don't believe that. I'd put some good money that if an excellent student is given an exact question from a previous year, they'll do better (faster & more accurate) on it, than when they're given a variation of it.
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#103Earlier quoted context omitted.
Not hardcoded, I think it's just likely that those problems exist in its training data in some form
I've always assumed they removed it, because it's such a basic and fundamental part of ML training that you separate your test and train data. And yet I never see any papers even mention if/how they do this. And I wonder if they do, how do they guarantee with high reliability that their massive terabytes of data don't contain the answer.
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#104The paper includes several examples of their modified questions. There has been a substantial jump from o1-preview to o1, so I gave several samples to o1 and o1-pro ( not o1-preview), and current o1s gave the correct answer to those modified problems. SOTA changes fast.
LLM boosters are so tiresome. You hardly did a rigorous evaluation, the set has been public since October and could have easily been added to the training data.
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#105Is it just an open secret that the models are currently just being hardcoded for random benchmarks? Seems weird that people would be asking Putnam problems to a chatbot :/
Not hardcoded, I think it's just likely that those problems exist in its training data in some form
"Solve (1503+5171)*(9494-4823)" reliably gets the correct answer from ChatGPT
"Write a poem about the solution to (1503+5171)*(9494-4823)" hallucinates an incorrect answer though
That suggests to me that they've papered over the models inability to do basic math, but it's a hack that doesn't generalize beyond the simplest cases.
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#106Earlier quoted context omitted.
Not hardcoded, I think it's just likely that those problems exist in its training data in some form
I've always assumed they removed it, because it's such a basic and fundamental part of ML training that you separate your test and train data. And yet I never see any papers even mention if/how they do this. And I wonder if they do, how do they guarantee with high reliability that their massive terabytes of data don't contain the answer.
Those are well known problems, that people talk about on different contexts. They would have to review their entire training set.
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#107Sure, LLMs will perform better if the answer to a problem is directly in their training set. But that doesn’t mean they perform bad when the answer isn’t in their training set.
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#108Performance of these LLMs on real life tasks feels very much like students last-minute cramming for Asian style exams. The ability to perfectly regurgitate, while no concept of meaning.
o3 is able to get 25% on never seen before frontiermath problems. sure, the models do better when the answer is directly in their dataset but they’ve already surpassed the average human in novelty on held out problems
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#109Earlier quoted context omitted.
Not hardcoded, I think it's just likely that those problems exist in its training data in some form
Seems a bit picky. If the bot has seen the exact problem before it's not really doing anything more than recall to solve it.
The problem is just that people keep insisting that those things are intelligent.
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#110Earlier quoted context omitted.
This is how patent disputes should be decided. If an LLM can figure it out, then it is not novel.
And what prompt would you give that does have novel input.