I remember when this stuff was all coming out and people were finally excited about ChatGPT getting the problem with "which is heavier, a 10 pound bag of feathers or a 10 pound bag of bricks?" problem correct. But of course it got it correct. It was in the training set. Vary the problem slightly by just changing the nouns, or changing the numbers so that one in fact was heavier than the other, and performance went al…
https://chatgpt.com/share/67756897-8974-8010-a0e0-c9e3b3e91f... so far o1-mini has bodied every task people are saying LLMs can’t do in this thread
30% drop in O1-preview accuracy when Putnam problems are slightly variated
211–220 of 558 posts
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#212Earlier quoted context omitted.
LLM boosters are so tiresome. You hardly did a rigorous evaluation, the set has been public since October and could have easily been added to the training data.
I'm not skilled enough in math to do a rigorous evaluation, so it was a quick check. Terence Tao is skilled enough, and he describes O1's math ability is "...roughly on par with a mediocre, but not completely incompetent graduate student" (good discussion at https://news.ycombinator.com/item?id=41540902 ), and the next iteration O3 just got 25% on his brand new Frontier Math test. Seeing LLMs as useless is banal, but…
Let it sink in how vague and almost meaningless that statement is.
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#213Earlier quoted context omitted.
https://chatgpt.com/share/67756897-8974-8010-a0e0-c9e3b3e91f... so far o1-mini has bodied every task people are saying LLMs can’t do in this thread
That appears to be the same model I used. This is why I emphasized I didn't "go shopping" for a result. That was the first result I got. I'm not at all surprised that it will nondeterministically get it correct sometimes. But if it doesn't get it correct every time, it doesn't "know". (In fact "going shopping" for errors would still even be fair. It should be correct all the time if it "knows". But it would be differ…
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#214I remember when this stuff was all coming out and people were finally excited about ChatGPT getting the problem with "which is heavier, a 10 pound bag of feathers or a 10 pound bag of bricks?" problem correct. But of course it got it correct. It was in the training set. Vary the problem slightly by just changing the nouns, or changing the numbers so that one in fact was heavier than the other, and performance went al…
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#215I remember when this stuff was all coming out and people were finally excited about ChatGPT getting the problem with "which is heavier, a 10 pound bag of feathers or a 10 pound bag of bricks?" problem correct. But of course it got it correct. It was in the training set. Vary the problem slightly by just changing the nouns, or changing the numbers so that one in fact was heavier than the other, and performance went al…
I got a very different answer:
A 10.01-pound bag of fluffy cotton is heavier than a 9.99-pound bag of steel ingots because 10.01 pounds is greater than 9.99 pounds. The material doesn't matter in this case; weight is the deciding factor.
What model returned your answer?
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#216Earlier quoted context omitted.
https://chatgpt.com/share/67756897-8974-8010-a0e0-c9e3b3e91f... so far o1-mini has bodied every task people are saying LLMs can’t do in this thread
This happens literally every time. Someone always says "ChatGPT can't do this!", but then when someone actually runs the example, chatGPT gets it right. Now what the OP is going to do next is proceed to move goalposts and say like "but umm I just asked chatgpt this, so clearly they modified the code in realtime to get the answer right"
that said, i don’t think this is a good test - i’ve seen it circling on twitter for months and it is almost certainly trained on similar tasks
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#217I remember when this stuff was all coming out and people were finally excited about ChatGPT getting the problem with "which is heavier, a 10 pound bag of feathers or a 10 pound bag of bricks?" problem correct. But of course it got it correct. It was in the training set. Vary the problem slightly by just changing the nouns, or changing the numbers so that one in fact was heavier than the other, and performance went al…
https://chatgpt.com/share/67756c29-111c-8002-b203-14c07ed1e6... I got a very different answer: A 10.01-pound bag of fluffy cotton is heavier than a 9.99-pound bag of steel ingots because 10.01 pounds is greater than 9.99 pounds. The material doesn't matter in this case; weight is the deciding factor. What model returned your answer?
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#218I remember when this stuff was all coming out and people were finally excited about ChatGPT getting the problem with "which is heavier, a 10 pound bag of feathers or a 10 pound bag of bricks?" problem correct. But of course it got it correct. It was in the training set. Vary the problem slightly by just changing the nouns, or changing the numbers so that one in fact was heavier than the other, and performance went al…
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#219The researcher's answer to their variant of "Year: 2016 ID: A1" in the appendix is wrong. The solution (sum of 1,2,5,6,9,10,13,14, ...) has an alternating pattern, so has to be two piecewise interleaved polynomials, which cannot be expressed as a single polyomial. Their answer works for k=1,2, but not k=3. https://openreview.net/pdf?id=YXnwlZe0yf This does not give me confidence in the results of their paper.
The statement doesn't hold for e.g. n=5. Taking m=2 gives the permutation (1 2 4 3), which is odd, and thus cannot have a square root.
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#220Earlier quoted context omitted.
o3 is able to get 25% on never seen before frontiermath problems. sure, the models do better when the answer is directly in their dataset but they’ve already surpassed the average human in novelty on held out problems
The average human did zero studying on representative problems. LLMs did a lot .
At top tier schools the most common score will usually be somewhere in the 0 to 10 range (out of a possible 120).