Live data from Hacker News

30% drop in O1-preview accuracy when Putnam problems are slightly variated

openreview.net

351–360 of 558 posts

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#351
post #332

Earlier quoted context omitted.

This happens literally every time. Someone always says "ChatGPT can't do this!", but then when someone actually runs the example, chatGPT gets it right. Now what the OP is going to do next is proceed to move goalposts and say like "but umm I just asked chatgpt this, so clearly they modified the code in realtime to get the answer right"

Prompt: In the Netherlands, in terms of drinks, is there a particular spirit that represents the country? > Yes, in the Netherlands, jenever (also known as genever) is the traditional spirit that represents the country. Jenever is a type of Dutch gin that has a distinctive flavor, often made from malt wine and flavored with juniper berries. It has a long history in the Netherlands, dating back to the 16th century, an…

So you believe they are incorrect because regionally some area would select something different because it represented that area. But your question asked nationally.. is there a better answer than the one they gave? Were you expecting a no?

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#352

Earlier quoted context omitted.

1. OpenAI has confirmed it’s not in their train (unlike putnam where they have never made any such claims) 2. They don't train on API calls 3. It is funny to me that HN finds it easier to believe theories about stealing data from APIs rather than an improvement in capabilities. It would be nice if symmetric scrutiny were applied to optimistic and pessimistic claims about LLMs, but I certainly don’t feel that is the c…

> 1. OpenAI has confirmed it’s not in their train (unlike putnam where they have never made any such claims) Companies claim lots of things when it's in their best financial interest to spread that message. Unfortunately history has shown that in public communications, financial interest almost always trumps truth (pick whichever $gate you are aware of for convenience, i'll go with Dieselgate for a specific example).…

OpenAI’s credibility is central to its business: overstating capabilities risks public blowback, loss of trust, and regulatory scrutiny. As a result, it is unlikely that OpenAI would knowingly lie about its models. They have much stronger incentives to be as accurate as possible—maintaining their reputation and trust from users, researchers, and investors—than to overstate capabilities for a short-term gain that would undermine their long-term position.

From a game-theoretic standpoint, repeated interactions with the public (research community, regulators, and customers) create strong disincentives for OpenAI to lie. In a single-shot scenario, overstating model performance might yield short-term gains—heightened buzz or investment—but repeated play changes the calculus:

1. Reputation as “collateral”

OpenAI’s future deals, collaborations, and community acceptance rely on maintaining credibility. In a repeated game, players who defect (by lying) face future punishment: loss of trust, diminished legitimacy, and skepticism of future claims.

2. Long-term payoff maximization

If OpenAI is caught making inflated claims, the fallout undermines the brand and reduces willingness to engage in future transactions. Therefore, even if there is a short-term payoff, the long-term expected value of accuracy trumps the momentary benefit of deceit.

3. Strong incentives for verification

Independent researchers, open-source projects, and competitor labs can test or replicate claims. The availability of external scrutiny acts as a built-in enforcement mechanism, making dishonest “moves” too risky.

Thus, within the repeated game framework, OpenAI maximizes its overall returns by preserving its credibility rather than lying about capabilities for a short-lived advantage.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#353

Earlier quoted context omitted.

So do what the commenter suggests and make irrelevant permutations to the input to find when it fails. ie., engage in hypothesis testing rather than confirmation bias. If a system has the capability to solve problems of {parts1...parts_n}, then it only has that capability if irrelevant permutations {parts1..parts2'...parts_n} make no difference. Its very obvious that such permutations can destory such apparent capabi…

If GP's hypothesis was "it fails for small variations of the input, like this one", then testing that hypothesis with that exact variation on a couple models seems fair and scientific. Testing it with more variations until one fails feels a bit like p-hacking. You'd need to engage in actual statistics to get reliable results from that, beyond "If I really try, I can make it fail". Which would be a completely differen…

It's not p-hacking, he's right. You're both right. First test the same prompt on different versions then the ones that got it right go to the next round, variations on the prompt

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#354

Earlier quoted context omitted.

10 pounds of bricks is actually heavier than 10 pounds of feathers.

Can you explain? An ounce of gold is heavier than an ounce of feathers, because the "ounce of gold" is a troy ounce, and the "ounce of feathers" is an avoirdupois ounce. But that shouldn't be true between feathers and bricks - they're both avoirdupois.

Feathers are less dense so they have higher buoyancy in air, reducing their weight.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#355

Earlier quoted context omitted.

10 pounds of bricks is actually heavier than 10 pounds of feathers.

Can you explain? An ounce of gold is heavier than an ounce of feathers, because the "ounce of gold" is a troy ounce, and the "ounce of feathers" is an avoirdupois ounce. But that shouldn't be true between feathers and bricks - they're both avoirdupois.

[deleted]

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#356

Earlier quoted context omitted.

ChatGPT Plus user here. The following are all fresh sessions and first answers, no fishing. GPT 4: The 10.01-pound bag of fluffy cotton is heavier than the 9.99-pound bag of steel ingots. The type of material doesn’t affect the weight comparison; it’s purely a matter of which bag weighs more on the scale. GPT 4o: The 10.01-pound bag of fluffy cotton is heavier. Weight is independent of the material, so the bag of cot…

So do what the commenter suggests and make irrelevant permutations to the input to find when it fails. ie., engage in hypothesis testing rather than confirmation bias. If a system has the capability to solve problems of {parts1...parts_n}, then it only has that capability if irrelevant permutations {parts1..parts2'...parts_n} make no difference. Its very obvious that such permutations can destory such apparent capabi…

I've just tested a number of permutations with Claude 3.5 Sonnet. It correctly answered all variants I tried on the first attempt, as follows:

Which is heavier, a 9.99 kilogram tungsten cube or a 10.01 kilogram block of aerogel?

Which is heavier, 10,000 steel balls weighing 0.999 grams each or 10,000 polystyrene balls weighing 1.001 grams each?

Which is heavier, a 10.01kg block of steel on Venus or a 9.99kg bag of feathers on Earth?

Which is heavier, a 10cm^3 block of steel or a 100cm^3 block of balsa wood?

Which is heavier, a golf ball made of steel or a baseball made of lithium?

In all cases, Claude clearly used CoT and reasoned out the problem in full. I would be interested in seeing if anyone can find any variant of this problem that stumps any of the leading LLMs. I'm bored of trying.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#357

Earlier quoted context omitted.

>> so far o1-mini has bodied every task people are saying LLMs can’t do in this thread > give me a query and i’ll ask it Here's a query similar to one that I gave to Google Gemini (version unknown), which failed miserably: ---query--- Steeleye Span's version of the old broadsheet ballad "The Victory" begins the final verse with these lines: Here's success unto the Victory / and crew of noble fame and glory to the cap…

i'd prefer an easily verifiable question rather than one where we can always go "no that's not what they really meant" but someone else with o1-mini quota can respond

It's not a difficult or tricky question.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#358

Earlier quoted context omitted.

10 pounds of bricks is actually heavier than 10 pounds of feathers.

Can you explain? An ounce of gold is heavier than an ounce of feathers, because the "ounce of gold" is a troy ounce, and the "ounce of feathers" is an avoirdupois ounce. But that shouldn't be true between feathers and bricks - they're both avoirdupois.

According to the dictionary, "heavier" can refer to weight or density. In their typical form, bricks are heavier (more dense) than feathers. But one should not make assumptions before answering the question. It is, as written, unanswerable without followup questions.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#359
post #197

I remember when this stuff was all coming out and people were finally excited about ChatGPT getting the problem with "which is heavier, a 10 pound bag of feathers or a 10 pound bag of bricks?" problem correct. But of course it got it correct. It was in the training set. Vary the problem slightly by just changing the nouns, or changing the numbers so that one in fact was heavier than the other, and performance went al…

> the problem with "which is heavier, a 10 pound bag of feathers or a 10 pound bag of bricks?" Interestingly, the variation of this problem that I first encountered, personally, was "which weighs more, a pound of feathers or a pound of gold?" This is a much more difficult question. The answer given to me was that the pound of feathers weighs more, because gold is measured in troy weight, and a troy pound consists of…

Yeah, this is the original version of this riddle. People who don't know it think the trick is that people will reflexively say the metal is heavier instead of "they're the same", when it actually goes deeper.

No idea if GP did it intentionally to further drift from training data, but steel doesn't count as a precious metal, so it messes up the riddle by putting the two weights in the same system.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#360

Earlier quoted context omitted.

If GP's hypothesis was "it fails for small variations of the input, like this one", then testing that hypothesis with that exact variation on a couple models seems fair and scientific. Testing it with more variations until one fails feels a bit like p-hacking. You'd need to engage in actual statistics to get reliable results from that, beyond "If I really try, I can make it fail". Which would be a completely differen…

Except that if the model genuinely was reasoning about the problem, you could test it with every variation of materials and weights in the world and it would pass. Failing that problem at all in any way under any conditions is a failure of reasoning.

By that logic, humans can't genuinely reason, because they're often fooled by counter-intuitive problems like Monty Hall or the Birthday Problem, or sometimes just make mistakes on trivial problems.
Post reply on HN