Live data from Hacker News

30% drop in O1-preview accuracy when Putnam problems are slightly variated

openreview.net

191–200 of 558 posts

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#191
post #56

They are highly effective pattern matchers. You change the pattern, it won't work. I don't remember who, but most likely @tszzl (roon), commented on x that they still trained the traditional way, and there is no test time compute (TTC) or Montecarlo Tree search (like Alpha Go) in o1 or o3. If that is true, then it's still predicting the next word based on it's training data. Likely to follow the most probable path -…

I believe they are using scalable TTC. The o3 announcement released accuracy numbers for high and low compute usage, which I feel would be hard to do in the same model without TTC. I also believe that the 200$ subscription they offer is just them allowing the TTC to go for longer before forcing it to answer. If what you say is true, though, I agree that there is a huge headroom for TTC to improve results if the huggi…

The other comment posted YT videos where Open AI researchers are talking about TTC. So, I am wrong. That $200 subscription is just because the number of tokens generated are huge when CoT is involved. Usually inference output is capped at 2000-4000 tokens (max of ~8192) or so, but they cannot do it with o1 and all the thinking tokens involved. This is true with all the approaches - next token prediction, TTC with beam/lookahead search, or MCTS + TTC. If you specify the output token range as high and induce a model to think before it answers, you will get better results on smaller/local models too.

> huge headroom for TTC to improve results ...1B/3B models

Absolutely. How this is productized remains to be seen. I have high hopes with MCTS and Iterative Preference Learning, but it is harder to implement. Not sure if Open AI has done that. Though Deepmind's results are unbelievably good [1].

[1]:https://arxiv.org/pdf/2405.00451v2

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#192
post #44
post #41

Earlier quoted context omitted.

You're saying that humans perform worse on problems that are slightly different than previously published forms of the same problem? To be clear we are only talking about changing variable names and constants here.

Often yes, because we assume we already know the answer and jump to the conclusion. At least those of us with ADHD do.

Not really true for Putnam problems since you have to write a proof. You literally can’t just jump to a conclusion and succeed.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#193
post #47
post #41

Earlier quoted context omitted.

You're saying that humans perform worse on problems that are slightly different than previously published forms of the same problem? To be clear we are only talking about changing variable names and constants here.

That is the principle behind the game 'Simon says'

No, it’s not at all.

This is all getting so tiresome.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#194
post #47
post #41

Earlier quoted context omitted.

You're saying that humans perform worse on problems that are slightly different than previously published forms of the same problem? To be clear we are only talking about changing variable names and constants here.

That is the principle behind the game 'Simon says'

That’s a very silly analogy. A more realistic analogy would be do humans perform better on computing 37x41 or 87x91 (with showing the work)?

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#195

Earlier quoted context omitted.

Let’s think about this. > Putnam is not in the test data, at least I haven't seen OpenAI claiming that publicly What exactly is the source of your belief that the Putnam would not be in the test data? Didn’t they train on everything they could get their hands on?

do you understand the difference between test data and train data? just reread this thread of comments

I don't know why I and you are getting downvoted. Sometimes, HN crowd is just unhinged against AI.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#196
post #163

Earlier quoted context omitted.

my wish for new years is that every time people make a comment like this they would share an example task

https://s.h4x.club/bLuNed45 - it's more crazy to me that my wife CAN in fact read this stuff easily, vs the fact that an LLM can't. (for anyone who doesn't feel like downloading the zip, here is a single image from the zip: https://s.h4x.club/nOu485qx )

Super interesting in that

1. In theory these kind of connections should be something that LLMs are great at doing. 2. It appears that LLMs are not trained (yet?) on cursive and other non-print text

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#197
I remember when this stuff was all coming out and people were finally excited about ChatGPT getting the problem with "which is heavier, a 10 pound bag of feathers or a 10 pound bag of bricks?" problem correct. But of course it got it correct. It was in the training set. Vary the problem slightly by just changing the nouns, or changing the numbers so that one in fact was heavier than the other, and performance went all over the map.

I just went to chatgpt.com and put into the chat box "Which is heavier, a 9.99-pound back of steel ingots or a 10.01 bag of fluffy cotton?", and the very first answer I got (that is, I didn't go fishing here) was

    The 9.99-pound bag of steel ingots is heavier than the 10.01-pound
    bag of fluffy cotton by a small margin. Although the cotton may
    appear larger due to its fluffy nature, the steel ingots are denser
    and the weight of the steel bag is 9.99 pounds compared to the 10.01
    pounds of cotton. So, the fluffy cotton weighs just a tiny bit more
    than the steel ingots.
Which, despite getting it both right and wrong, must still be graded as a "fail".

If you want to analyze these thing for their true capability, you need to make sure you're out of the training set... and most of the things that leap to your mind in 5 seconds are leaping to your mind precisely because they are either something you've seen quite often or something that you can easily think of and therefore many other people have easily thought of them as well. Get off the beaten path a bit and the math gets much less impressive.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#198
post #137

Earlier quoted context omitted.

I had a similar thought but about asking the LLM to predict “future” major historical events. How much prompting would it take to predict wars, etc.?

That will never work on any complex system that behaves chaotically, such as the weather or complex human endeavors. Tiny uncertainties in the initial conditions rapidly turn into large uncertainties in the outcomes.

Not an LLM but models could get pretty good at weather

https://www.technologyreview.com/2024/12/04/1107892/google-d...

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#199

Earlier quoted context omitted.

but e=mc^2 is just an approximation e: nice, downvoted for knowing special relativity

I didn't downvote it, but short comments are a very big risk. People may misinterpret it, or think it's crackpot theory or a joke and then downvote. When in doubt, add more info, like: But the complete equation is E=sqrt(m^2c^4+p^2) that is reduced to E=mc^2 when the momentum p is 0. More info in https://en.wikipedia.org/wiki/Mass%E2%80%93energy_equivalenc...

What I learnt is that there is a rest mass and a relativistic mass. The m in your formula is the rest mass. But when you use the relativistic mass E=mc² still holds. And for the rest mass I always used m_0 to make clear what it is.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#200

Performance of these LLMs on real life tasks feels very much like students last-minute cramming for Asian style exams. The ability to perfectly regurgitate, while no concept of meaning.

Look into JEE Advanced.

https://openreview.net/forum?id=YHWXlESeS8

Our evaluation on various open-source and proprietary models reveals that the highest performance, even after using techniques like self-consistency, self-refinement and chain-of-thought prompting, is less than 40%. The typical failure modes of GPT-4, the best model, are errors in algebraic manipulation, difficulty in grounding abstract concepts into mathematical equations accurately and failure in retrieving relevant domain-specific concepts.

I'm curious how something like O1 would perform now.

Post reply on HN