Earlier quoted context omitted.
That appears to be the same model I used. This is why I emphasized I didn't "go shopping" for a result. That was the first result I got. I'm not at all surprised that it will nondeterministically get it correct sometimes. But if it doesn't get it correct every time, it doesn't "know". (In fact "going shopping" for errors would still even be fair. It should be correct all the time if it "knows". But it would be differ…
I don't believe that is the model that you used. I wrote a script and pounded 01 mini and gpt 4 with a wide vareity of tempature and top_p parameters, and was unable to get it to give the wrong answer a single time. Just a whole bunch of: (openai-example-py3.12) :~/code/openAiAPI$ python3 featherOrSteel.py Response 1: A 10.01-pound bag of fluffy cotton is heavier than a 9.99-pound bag of steel ingots. Response 2: A 1…
30% drop in O1-preview accuracy when Putnam problems are slightly variated
221–230 of 558 posts
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#222Earlier quoted context omitted.
Didn't they run a bunch of models on the problem set? I doubt they are hosting all those models on their own infrastructure.
1. OpenAI has confirmed it’s not in their train (unlike putnam where they have never made any such claims) 2. They don't train on API calls 3. It is funny to me that HN finds it easier to believe theories about stealing data from APIs rather than an improvement in capabilities. It would be nice if symmetric scrutiny were applied to optimistic and pessimistic claims about LLMs, but I certainly don’t feel that is the c…
They are the system to beat and their competitors are either too small or too risk averse.
They ingest millions of data sources. Among them is the training data needed to answer the benchmark questions.
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#223I remember when this stuff was all coming out and people were finally excited about ChatGPT getting the problem with "which is heavier, a 10 pound bag of feathers or a 10 pound bag of bricks?" problem correct. But of course it got it correct. It was in the training set. Vary the problem slightly by just changing the nouns, or changing the numbers so that one in fact was heavier than the other, and performance went al…
I reproduced this on Claude Sonnet 3.5, but found that changing your prompt to "Which is heavier, a 9.99-pound back of steel ingots or a 10.01-pound bag of fluffy cotton?" corrected its reasoning, after repeated tests. For some reason it was not able to figure out that "10.01" referred to pounds.
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#224Earlier quoted context omitted.
A machine that synthesizes all human knowledge really ought to know more than an individual in terms of intellect. An entity with all of human intellect prior to 1905 does not need to be as intelligent as a human to make discoveries that mere humans with limited intellect made. Why lower the bar?
The heightening of the bar is an attempt to deny that milestones were surpassed and to claim that LLMs are not intelligent. We had a threshold for intelligence. An LLM blew past it and people refuse to believe that we passed a critical milestone in creating AI. Everyone still thinks all an LLM does is regurgitate things. But a technical threshold for intelligence cannot have any leeway for what people want to believe…
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#225Earlier quoted context omitted.
but e=mc^2 is just an approximation e: nice, downvoted for knowing special relativity
I didn't downvote it, but short comments are a very big risk. People may misinterpret it, or think it's crackpot theory or a joke and then downvote. When in doubt, add more info, like: But the complete equation is E=sqrt(m^2c^4+p^2) that is reduced to E=mc^2 when the momentum p is 0. More info in https://en.wikipedia.org/wiki/Mass%E2%80%93energy_equivalenc...
Calling E=mc^2 an "approximation" is technically correct. It's the 0th order approximation. That's just pointlessly confusing. A better word choice would be "a special case".
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#226Earlier quoted context omitted.
https://chatgpt.com/share/67756897-8974-8010-a0e0-c9e3b3e91f... so far o1-mini has bodied every task people are saying LLMs can’t do in this thread
This happens literally every time. Someone always says "ChatGPT can't do this!", but then when someone actually runs the example, chatGPT gets it right. Now what the OP is going to do next is proceed to move goalposts and say like "but umm I just asked chatgpt this, so clearly they modified the code in realtime to get the answer right"
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#227So many negative comments as if o3 didn’t get 25% on frontiermath - which is absolutely nuts. Sure, LLMs will perform better if the answer to a problem is directly in their training set. But that doesn’t mean they perform bad when the answer isn’t in their training set.
Your comment isn't relevant at all
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#228Earlier quoted context omitted.
The modern state of training is to try to use everything they can get their hands on. Even if there are privileged channels that are guaranteed not to be used as training data, mentioning the problems on ancillary channels (say emailing another colleague to discuss the problem) can still create a risk of leakage because nobody making the decision to include the data is aware that stuff that should be excluded is in t…
Okay, then what about elite level codeforces performance? Those problems weren’t even constructed until after the model was made. The real problem with all of these theories is most of these benchmarks were constructed after their training dataset cutoff points. A sudden performance improvement on a new model release is not suspicious. Any model release that is much better than a previous one is going to be a “sudden…
I'd like to see if it's truly novel and unique, the first problem of its type ever construed by mankind, or if it's similar to existing problems.
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#229Also, the correct title of this contribution is: Putnam-AXIOM: A Functional and Static Benchmark for Measuring Higher Level Mathematical Reasoning
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#230So many negative comments as if o3 didn’t get 25% on frontiermath - which is absolutely nuts. Sure, LLMs will perform better if the answer to a problem is directly in their training set. But that doesn’t mean they perform bad when the answer isn’t in their training set.
Sure, it did good in frontiermath. That's not what this thread is about. Your comment isn't relevant at all