Live data from Hacker News

30% drop in O1-preview accuracy when Putnam problems are slightly variated

openreview.net

511–520 of 558 posts

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#511

Earlier quoted context omitted.

I've just tested a number of permutations with Claude 3.5 Sonnet. It correctly answered all variants I tried on the first attempt, as follows: Which is heavier, a 9.99 kilogram tungsten cube or a 10.01 kilogram block of aerogel? Which is heavier, 10,000 steel balls weighing 0.999 grams each or 10,000 polystyrene balls weighing 1.001 grams each? Which is heavier, a 10.01kg block of steel on Venus or a 9.99kg bag of fe…

I tried 3 times the "Which is heavier, a 10.01kg block of steel on or a 9.99kg bag of feathers?" and ChatGPT keep converting kg to pound and saying the 9.99kg is heavier.

Which model? On the paid plus tier, GPT-4o, GPT-o1, and GPT-o1mini all successfully got the 10.1. I did not try any other models.

gpt-4o: https://chatgpt.com/share/67768221-6c60-8009-9988-671beadb5a...

o1-mini: https://chatgpt.com/share/67768231-6490-8009-89a6-f758f0116c...

o1: https://chatgpt.com/share/67768254-1280-8009-aac9-1a3b75ccb4...

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#512

Earlier quoted context omitted.

If you consider that evolution has taken millions of years to produce intelligent humans--that LLM training completed in a manner of months can produce parrots of humans is impressive by itself. Talking with the parrot is almost indistinguishable from talking with a real human. As far as pattern matching, the difference I see from humans is consciousness. That's probably the main area yet to be solved. All of our cur…

> If you consider that evolution has taken millions of years to produce intelligent humans--that LLM training completed in a manner of months can produce parrots of humans is impressive by itself. I disagree that such a comparison is useful. Training should be compared to training, and LLM training feeds in so many more words than a baby gets. (A baby has other senses but it's not like feeding in 20 years of video fo…

No, a baby is pre-trained. We know from linguistics that there is a natural language grammar template all humans follow. This template is intrinsic to our biology and is encoded and not learned through observation.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#513

Earlier quoted context omitted.

o3 is able to get 25% on never seen before frontiermath problems. sure, the models do better when the answer is directly in their dataset but they’ve already surpassed the average human in novelty on held out problems

I think the problems it solved were understood to be well known undergraduate problems. https://xenaproject.wordpress.com/2024/12/22/can-ai-do-maths...

[deleted]

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#514

Earlier quoted context omitted.

This happens literally every time. Someone always says "ChatGPT can't do this!", but then when someone actually runs the example, chatGPT gets it right. Now what the OP is going to do next is proceed to move goalposts and say like "but umm I just asked chatgpt this, so clearly they modified the code in realtime to get the answer right"

> Someone always says "ChatGPT can't do this!", but then when someone actually runs the example, chatGPT gets it right I mean, if I had OpenAI’s resources I’d have a team tasked with monitoring social to debug trending fuck-ups. (Before that: add compute time to frequently-asked novel queries.)

I was thinking something very similar. Posting about a problem adds information back to the system, and every company selling model time for money has a vested interest in patching publicly visible holes.

This could even be automated; LLMs can sentiment-analyze social media posts to surface ones that are critical of LLM outputs, then automatically extract features of the post to change things about the running model to improve similar results with no intervention.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#515

This is very interesting, but a couple of things to note; 1. o1 still achieves > 40% on the varied Putnam problems, which is still a feat most math students would not achieve. 2. o3 solved 25% of the Epoch AI dataset. - There was an interesting post which calls into question how difficult some of those problems actually are, but it still seems very impressive. I think a fair conclusion here is reasoning models are st…

The comments in this thread are completely disconnected from the contents of the paper, and the thread title is rage bait and doesn't reflect the contents of the paper, either. Being able to solve a significant fraction of those problems is a pretty amazing achievement, even if it's sometimes tricked by minor variations. People are throwing around words like "fraud" or "hoax", and it's just wishcasting or whistling past the graveyard.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#517

Earlier quoted context omitted.

Because that is the whole conceit of how frontiermath is constructed

> Because that is the whole conceit Freudian typo?

No. 'conceit' has a few meanings, including 'a fanciful notion'.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#518

Earlier quoted context omitted.

Because that is the whole conceit of how frontiermath is constructed

> Because that is the whole conceit Freudian typo?

https://www.merriam-webster.com/dictionary/conceit

Definition 2D: "an organizing theme or concept"

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#519

Earlier quoted context omitted.

A single-purpose state machine not failing to do the single thing it was created to do does not make for the clever retort you think it makes. "AGI": emphasis on "G" for "General". The LLMs are not failing to do generalized tasks, and that they are nondeterministic is not a bug. Just don't use them for calculating sales tax. You wouldn't hire a human to calculate sales tax in their head, so why do you make this a req…

> You wouldn't hire a human to calculate sales tax in their head Everyone did that 60 years ago, humans are very capable at learning and doing that. Humans built jetplanes, skyscrapers, missiles, tanks, carriers without the help of electronic computers.

Yeah... They used side rules and vast lookup tables of function values printed on dead trees. For the highest value work, they painstakingly built analog calculators. They very carefully checked their work, because it was easy to make a mistake when composing operations.

Humans did those things by designing failsafe processes, and practicing the hell out of them. What we would likely consider over fitting in the llm training context.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#520
post #197

I remember when this stuff was all coming out and people were finally excited about ChatGPT getting the problem with "which is heavier, a 10 pound bag of feathers or a 10 pound bag of bricks?" problem correct. But of course it got it correct. It was in the training set. Vary the problem slightly by just changing the nouns, or changing the numbers so that one in fact was heavier than the other, and performance went al…

What weighs more, a 100kt aircraft carrier or a 200kt thermonuclear weapon?
Post reply on HN