Live data from Hacker News

30% drop in O1-preview accuracy when Putnam problems are slightly variated

openreview.net

471–480 of 558 posts

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#471

Or it's time to step back and call it what it is - very good pattern recognition. I mean, that's cool... we can get a lot of work done with pattern recognition. Most of the human race never really moves above that level of thinking in the workforce or navigating their daily life, especially if they default to various societally prescribed patterns of getting stuff done (eg. go to college or the military , find a job…

> Or it's time to step back and call it what it is - very good pattern recognition. Or maybe it's time to stop wheeling out this tedious and disingenuous dismissal. Saying it is just "pattern recognition" (or a "stochastic parrot") implies behavioural and performance characteristics that have very clearly been greatly exceeded.

so how much you have riding on nvidia bro?

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#472
post #197

I remember when this stuff was all coming out and people were finally excited about ChatGPT getting the problem with "which is heavier, a 10 pound bag of feathers or a 10 pound bag of bricks?" problem correct. But of course it got it correct. It was in the training set. Vary the problem slightly by just changing the nouns, or changing the numbers so that one in fact was heavier than the other, and performance went al…

fwiw i think reasoning models have at least solved this. even the smallest reasoning model, o1-mini, gets it right first try on my test:

  A 10.01-pound bag of fluffy cotton is heavier than a 9.99-pound bag of steel ingots. Despite the significant difference in density and volume between steel and cotton, the weights provided clearly indicate that the cotton bag has a greater mass.

  Summary:

  Steel ingots: 9.99 pounds
  Fluffy cotton: 10.01 pounds
  Conclusion: The 10.01-pound bag of cotton is heavier.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#473

Earlier quoted context omitted.

frankly i’m not sure what standard you would possibly consider a debunking codeforces constantly adds new problems that’s like the entire point of the contest, no?

OpenAI never actually directly participated in the contests. OpenAI says they took (unspecified) problems and "simulated" the score it would have had. If they solved recent contests in a realistic contest simulation I would expect them to give the actual solutions and success rates as well, like they did for IOI problems, so I'm actually confused as to why they didn't.

very good clarification, thanks - they should absolutely release more details to provide more clarity and ideally just participate live? i suspect that the model takes a while for individual problems so time might be a constraint there

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#474
post #385
post #345

Earlier quoted context omitted.

It seems more like ChatGPT was asked a rather bizarre question with far too little detail to make sense, and ChatGPT failed to notice or to ask for more information. Although it did get rather impressively confused about the pressure of the air.

“Failed to notice” is a hilariously weird way of talking about tokenization.

Tokenization?

I mean that ChatGPT had no questions about the container of the gas (does it have weight? is it tared out?) or about buoyancy. And it’s really rather sad that ChatGPT calculated the volume of air at atmospheric temperature and pressure and didn’t notice that it was nowhere near fitting into its supposed container. (At least 1.01lb of liquid nitrogen would fit.)

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#475

Earlier quoted context omitted.

I've just tested a number of permutations with Claude 3.5 Sonnet. It correctly answered all variants I tried on the first attempt, as follows: Which is heavier, a 9.99 kilogram tungsten cube or a 10.01 kilogram block of aerogel? Which is heavier, 10,000 steel balls weighing 0.999 grams each or 10,000 polystyrene balls weighing 1.001 grams each? Which is heavier, a 10.01kg block of steel on Venus or a 9.99kg bag of fe…

BTW - the model may be wrong depending on the example. More voluminous objects displace more air and due to buoyancy are lighter for the same mass. The proper way to ask it would be to ask which object has more mass.

For what it’s worth, I asked it the following:

Q:

Which weighs more on earth: 100kg of steel or a 100kg polystyrene ball? Be as technically correct as possible

A: Technically, 100 kg of steel and 100 kg of polystyrene weigh the same on Earth when measured on a scale, as they both experience the same gravitational force due to their equal mass. However, there are some nuances to consider:

Buoyancy Effect

Earth’s atmosphere exerts an upward buoyant force on objects, proportional to the volume of the object and the density of the surrounding air. Since polystyrene has a much lower density than steel, a 100 kg polystyrene ball will have a significantly larger volume and experience a greater buoyant force compared to 100 kg of steel.

As a result: • Effective weight (apparent weight): The steel will weigh slightly more than the polystyrene on a scale because the buoyant force acting on the polystyrene reduces its apparent weight more.

Actual Weight

Weight is the gravitational force on an object and is given by:

Where and . Without considering buoyancy, both objects have the same weight of approximately (newtons).

Conclusion • True weight (in a vacuum): Both are identical at 981 N. • Apparent weight (in air): Steel is slightly heavier due to reduced buoyant force acting on it compared to the polystyrene ball.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#476
post #444

Earlier quoted context omitted.

This definition alone might be fine enough if the word "intelligence" wasn't already widely used outside of AI research. It is though, and the idea that intelligence is measured solely through economic value is a very, very strange approach. Try applying that definition to humans and you pretty quickly run into issues, both moral and practical. It also invalidates basically anything we've done over centuries consider…

> This definition alone might be fine enough if the word "intelligence" wasn't already widely used outside of AI research. It is though, and the idea that intelligence is measured solely through economic value is a very, very strange approach. The response from @s1mplicissimus' on my previous comment is asking about "common usage" definitions of intelligence, and this is (IMO unfortunately) one of the many "common us…

> smart people generally earn more.

> there's been a widespread belief that being wealthy is the proof of superiority.

Both of these are assumptions though, and working in the reverse order. Its one thing to expect that intelligence will lead to higher value outcomes and entirely different to expect that higher value outcomes prove intelligence.

It seems reasonable that higher intelligence, combined with the incentives if a capitalist system, will lead to higher intelligence people getting more wealthy. They learn to play the game and find ways to "win."

It seems unreasonable to assume that anyone or anything that "wins" in that system much be more intelligent. Said differently, intelligence may lead to wealth but wealth doesn't imply intelligence.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#477
post #165

Earlier quoted context omitted.

but e=mc^2 is just an approximation e: nice, downvoted for knowing special relativity

Can you elaborate? How is E=mc^2 an approximation, in special relativity or otherwise? What is it an approximation of?

E^2 = (mc^2)^2 + (pc)^2,

where p is momentum. When an object is traveling at relativistic speeds, the momentum forms a more significant portion of its energy

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#478

Earlier quoted context omitted.

I've read the thread and I think it's not very coherent overall, I also not sure if we disagree =) I agree that having putnam problems on OpenAI training set is not a smoking gun, however it's (almost) certain they are on training set, and having them would affect performance of the model on them too. Hence research like this is important, since it shows that observed behavior of the models is memoization to large ex…

nobody serious (like OAI) was using the putnam problems to claim generalization. this is a refutation in search of a claim - and many people in the upstream thread are suggesting that OAI is doing something wrong by training on a benchmark. OAI uses datasets like frontiermath or arc-agi that are actually held out to evaluate generalization.

I, actually, would disagree with this. To me ability to solve frontiermath does imply ability to solve putnam problems too. Only with putnam problems being easier - they are already been seen by the model, and they are also simpler problems. And just like this - putnam problem with simple changes are also one of the easier stops on the way to truly generalizing math models, with frontiermath being one of the last stops on the way there.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#479

Earlier quoted context omitted.

So do what the commenter suggests and make irrelevant permutations to the input to find when it fails. ie., engage in hypothesis testing rather than confirmation bias. If a system has the capability to solve problems of {parts1...parts_n}, then it only has that capability if irrelevant permutations {parts1..parts2'...parts_n} make no difference. Its very obvious that such permutations can destory such apparent capabi…

I've just tested a number of permutations with Claude 3.5 Sonnet. It correctly answered all variants I tried on the first attempt, as follows: Which is heavier, a 9.99 kilogram tungsten cube or a 10.01 kilogram block of aerogel? Which is heavier, 10,000 steel balls weighing 0.999 grams each or 10,000 polystyrene balls weighing 1.001 grams each? Which is heavier, a 10.01kg block of steel on Venus or a 9.99kg bag of fe…

I tried 3 times the "Which is heavier, a 10.01kg block of steel on or a 9.99kg bag of feathers?" and ChatGPT keep converting kg to pound and saying the 9.99kg is heavier.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#480
post #204

Earlier quoted context omitted.

That appears to be the same model I used. This is why I emphasized I didn't "go shopping" for a result. That was the first result I got. I'm not at all surprised that it will nondeterministically get it correct sometimes. But if it doesn't get it correct every time, it doesn't "know". (In fact "going shopping" for errors would still even be fair. It should be correct all the time if it "knows". But it would be differ…

I don't believe that is the model that you used. I wrote a script and pounded 01 mini and gpt 4 with a wide vareity of tempature and top_p parameters, and was unable to get it to give the wrong answer a single time. Just a whole bunch of: (openai-example-py3.12) :~/code/openAiAPI$ python3 featherOrSteel.py Response 1: A 10.01-pound bag of fluffy cotton is heavier than a 9.99-pound bag of steel ingots. Response 2: A 1…

The elephant in the room is that HN is full of people facing an existential threat.
Post reply on HN