Earlier quoted context omitted.
This happens literally every time. Someone always says "ChatGPT can't do this!", but then when someone actually runs the example, chatGPT gets it right. Now what the OP is going to do next is proceed to move goalposts and say like "but umm I just asked chatgpt this, so clearly they modified the code in realtime to get the answer right"
Prompt: In the Netherlands, in terms of drinks, is there a particular spirit that represents the country? > Yes, in the Netherlands, jenever (also known as genever) is the traditional spirit that represents the country. Jenever is a type of Dutch gin that has a distinctive flavor, often made from malt wine and flavored with juniper berries. It has a long history in the Netherlands, dating back to the 16th century, an…
30% drop in O1-preview accuracy when Putnam problems are slightly variated
351–360 of 558 posts
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#352Earlier quoted context omitted.
1. OpenAI has confirmed it’s not in their train (unlike putnam where they have never made any such claims) 2. They don't train on API calls 3. It is funny to me that HN finds it easier to believe theories about stealing data from APIs rather than an improvement in capabilities. It would be nice if symmetric scrutiny were applied to optimistic and pessimistic claims about LLMs, but I certainly don’t feel that is the c…
> 1. OpenAI has confirmed it’s not in their train (unlike putnam where they have never made any such claims) Companies claim lots of things when it's in their best financial interest to spread that message. Unfortunately history has shown that in public communications, financial interest almost always trumps truth (pick whichever $gate you are aware of for convenience, i'll go with Dieselgate for a specific example).…
From a game-theoretic standpoint, repeated interactions with the public (research community, regulators, and customers) create strong disincentives for OpenAI to lie. In a single-shot scenario, overstating model performance might yield short-term gains—heightened buzz or investment—but repeated play changes the calculus:
1. Reputation as “collateral”
OpenAI’s future deals, collaborations, and community acceptance rely on maintaining credibility. In a repeated game, players who defect (by lying) face future punishment: loss of trust, diminished legitimacy, and skepticism of future claims.
2. Long-term payoff maximization
If OpenAI is caught making inflated claims, the fallout undermines the brand and reduces willingness to engage in future transactions. Therefore, even if there is a short-term payoff, the long-term expected value of accuracy trumps the momentary benefit of deceit.
3. Strong incentives for verification
Independent researchers, open-source projects, and competitor labs can test or replicate claims. The availability of external scrutiny acts as a built-in enforcement mechanism, making dishonest “moves” too risky.
Thus, within the repeated game framework, OpenAI maximizes its overall returns by preserving its credibility rather than lying about capabilities for a short-lived advantage.
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#353Earlier quoted context omitted.
So do what the commenter suggests and make irrelevant permutations to the input to find when it fails. ie., engage in hypothesis testing rather than confirmation bias. If a system has the capability to solve problems of {parts1...parts_n}, then it only has that capability if irrelevant permutations {parts1..parts2'...parts_n} make no difference. Its very obvious that such permutations can destory such apparent capabi…
If GP's hypothesis was "it fails for small variations of the input, like this one", then testing that hypothesis with that exact variation on a couple models seems fair and scientific. Testing it with more variations until one fails feels a bit like p-hacking. You'd need to engage in actual statistics to get reliable results from that, beyond "If I really try, I can make it fail". Which would be a completely differen…
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#354Earlier quoted context omitted.
10 pounds of bricks is actually heavier than 10 pounds of feathers.
Can you explain? An ounce of gold is heavier than an ounce of feathers, because the "ounce of gold" is a troy ounce, and the "ounce of feathers" is an avoirdupois ounce. But that shouldn't be true between feathers and bricks - they're both avoirdupois.
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#355Earlier quoted context omitted.
10 pounds of bricks is actually heavier than 10 pounds of feathers.
Can you explain? An ounce of gold is heavier than an ounce of feathers, because the "ounce of gold" is a troy ounce, and the "ounce of feathers" is an avoirdupois ounce. But that shouldn't be true between feathers and bricks - they're both avoirdupois.
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#356Earlier quoted context omitted.
ChatGPT Plus user here. The following are all fresh sessions and first answers, no fishing. GPT 4: The 10.01-pound bag of fluffy cotton is heavier than the 9.99-pound bag of steel ingots. The type of material doesn’t affect the weight comparison; it’s purely a matter of which bag weighs more on the scale. GPT 4o: The 10.01-pound bag of fluffy cotton is heavier. Weight is independent of the material, so the bag of cot…
So do what the commenter suggests and make irrelevant permutations to the input to find when it fails. ie., engage in hypothesis testing rather than confirmation bias. If a system has the capability to solve problems of {parts1...parts_n}, then it only has that capability if irrelevant permutations {parts1..parts2'...parts_n} make no difference. Its very obvious that such permutations can destory such apparent capabi…
Which is heavier, a 9.99 kilogram tungsten cube or a 10.01 kilogram block of aerogel?
Which is heavier, 10,000 steel balls weighing 0.999 grams each or 10,000 polystyrene balls weighing 1.001 grams each?
Which is heavier, a 10.01kg block of steel on Venus or a 9.99kg bag of feathers on Earth?
Which is heavier, a 10cm^3 block of steel or a 100cm^3 block of balsa wood?
Which is heavier, a golf ball made of steel or a baseball made of lithium?
In all cases, Claude clearly used CoT and reasoned out the problem in full. I would be interested in seeing if anyone can find any variant of this problem that stumps any of the leading LLMs. I'm bored of trying.
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#357Earlier quoted context omitted.
>> so far o1-mini has bodied every task people are saying LLMs can’t do in this thread > give me a query and i’ll ask it Here's a query similar to one that I gave to Google Gemini (version unknown), which failed miserably: ---query--- Steeleye Span's version of the old broadsheet ballad "The Victory" begins the final verse with these lines: Here's success unto the Victory / and crew of noble fame and glory to the cap…
i'd prefer an easily verifiable question rather than one where we can always go "no that's not what they really meant" but someone else with o1-mini quota can respond
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#358Earlier quoted context omitted.
10 pounds of bricks is actually heavier than 10 pounds of feathers.
Can you explain? An ounce of gold is heavier than an ounce of feathers, because the "ounce of gold" is a troy ounce, and the "ounce of feathers" is an avoirdupois ounce. But that shouldn't be true between feathers and bricks - they're both avoirdupois.
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#359I remember when this stuff was all coming out and people were finally excited about ChatGPT getting the problem with "which is heavier, a 10 pound bag of feathers or a 10 pound bag of bricks?" problem correct. But of course it got it correct. It was in the training set. Vary the problem slightly by just changing the nouns, or changing the numbers so that one in fact was heavier than the other, and performance went al…
> the problem with "which is heavier, a 10 pound bag of feathers or a 10 pound bag of bricks?" Interestingly, the variation of this problem that I first encountered, personally, was "which weighs more, a pound of feathers or a pound of gold?" This is a much more difficult question. The answer given to me was that the pound of feathers weighs more, because gold is measured in troy weight, and a troy pound consists of…
No idea if GP did it intentionally to further drift from training data, but steel doesn't count as a precious metal, so it messes up the riddle by putting the two weights in the same system.
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#360Earlier quoted context omitted.
If GP's hypothesis was "it fails for small variations of the input, like this one", then testing that hypothesis with that exact variation on a couple models seems fair and scientific. Testing it with more variations until one fails feels a bit like p-hacking. You'd need to engage in actual statistics to get reliable results from that, beyond "If I really try, I can make it fail". Which would be a completely differen…
Except that if the model genuinely was reasoning about the problem, you could test it with every variation of materials and weights in the world and it would pass. Failing that problem at all in any way under any conditions is a failure of reasoning.