Or it's time to step back and call it what it is - very good pattern recognition. I mean, that's cool... we can get a lot of work done with pattern recognition. Most of the human race never really moves above that level of thinking in the workforce or navigating their daily life, especially if they default to various societally prescribed patterns of getting stuff done (eg. go to college or the military , find a job…
> Or it's time to step back and call it what it is - very good pattern recognition. Or maybe it's time to stop wheeling out this tedious and disingenuous dismissal. Saying it is just "pattern recognition" (or a "stochastic parrot") implies behavioural and performance characteristics that have very clearly been greatly exceeded.
30% drop in O1-preview accuracy when Putnam problems are slightly variated
471–480 of 558 posts
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#472I remember when this stuff was all coming out and people were finally excited about ChatGPT getting the problem with "which is heavier, a 10 pound bag of feathers or a 10 pound bag of bricks?" problem correct. But of course it got it correct. It was in the training set. Vary the problem slightly by just changing the nouns, or changing the numbers so that one in fact was heavier than the other, and performance went al…
A 10.01-pound bag of fluffy cotton is heavier than a 9.99-pound bag of steel ingots. Despite the significant difference in density and volume between steel and cotton, the weights provided clearly indicate that the cotton bag has a greater mass.
Summary:
Steel ingots: 9.99 pounds
Fluffy cotton: 10.01 pounds
Conclusion: The 10.01-pound bag of cotton is heavier.Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#473Earlier quoted context omitted.
frankly i’m not sure what standard you would possibly consider a debunking codeforces constantly adds new problems that’s like the entire point of the contest, no?
OpenAI never actually directly participated in the contests. OpenAI says they took (unspecified) problems and "simulated" the score it would have had. If they solved recent contests in a realistic contest simulation I would expect them to give the actual solutions and success rates as well, like they did for IOI problems, so I'm actually confused as to why they didn't.
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#474Earlier quoted context omitted.
It seems more like ChatGPT was asked a rather bizarre question with far too little detail to make sense, and ChatGPT failed to notice or to ask for more information. Although it did get rather impressively confused about the pressure of the air.
“Failed to notice” is a hilariously weird way of talking about tokenization.
I mean that ChatGPT had no questions about the container of the gas (does it have weight? is it tared out?) or about buoyancy. And it’s really rather sad that ChatGPT calculated the volume of air at atmospheric temperature and pressure and didn’t notice that it was nowhere near fitting into its supposed container. (At least 1.01lb of liquid nitrogen would fit.)
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#475Earlier quoted context omitted.
I've just tested a number of permutations with Claude 3.5 Sonnet. It correctly answered all variants I tried on the first attempt, as follows: Which is heavier, a 9.99 kilogram tungsten cube or a 10.01 kilogram block of aerogel? Which is heavier, 10,000 steel balls weighing 0.999 grams each or 10,000 polystyrene balls weighing 1.001 grams each? Which is heavier, a 10.01kg block of steel on Venus or a 9.99kg bag of fe…
BTW - the model may be wrong depending on the example. More voluminous objects displace more air and due to buoyancy are lighter for the same mass. The proper way to ask it would be to ask which object has more mass.
Q:
Which weighs more on earth: 100kg of steel or a 100kg polystyrene ball? Be as technically correct as possible
A: Technically, 100 kg of steel and 100 kg of polystyrene weigh the same on Earth when measured on a scale, as they both experience the same gravitational force due to their equal mass. However, there are some nuances to consider:
Buoyancy Effect
Earth’s atmosphere exerts an upward buoyant force on objects, proportional to the volume of the object and the density of the surrounding air. Since polystyrene has a much lower density than steel, a 100 kg polystyrene ball will have a significantly larger volume and experience a greater buoyant force compared to 100 kg of steel.
As a result: • Effective weight (apparent weight): The steel will weigh slightly more than the polystyrene on a scale because the buoyant force acting on the polystyrene reduces its apparent weight more.
Actual Weight
Weight is the gravitational force on an object and is given by:
Where and . Without considering buoyancy, both objects have the same weight of approximately (newtons).
Conclusion • True weight (in a vacuum): Both are identical at 981 N. • Apparent weight (in air): Steel is slightly heavier due to reduced buoyant force acting on it compared to the polystyrene ball.
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#476Earlier quoted context omitted.
This definition alone might be fine enough if the word "intelligence" wasn't already widely used outside of AI research. It is though, and the idea that intelligence is measured solely through economic value is a very, very strange approach. Try applying that definition to humans and you pretty quickly run into issues, both moral and practical. It also invalidates basically anything we've done over centuries consider…
> This definition alone might be fine enough if the word "intelligence" wasn't already widely used outside of AI research. It is though, and the idea that intelligence is measured solely through economic value is a very, very strange approach. The response from @s1mplicissimus' on my previous comment is asking about "common usage" definitions of intelligence, and this is (IMO unfortunately) one of the many "common us…
> there's been a widespread belief that being wealthy is the proof of superiority.
Both of these are assumptions though, and working in the reverse order. Its one thing to expect that intelligence will lead to higher value outcomes and entirely different to expect that higher value outcomes prove intelligence.
It seems reasonable that higher intelligence, combined with the incentives if a capitalist system, will lead to higher intelligence people getting more wealthy. They learn to play the game and find ways to "win."
It seems unreasonable to assume that anyone or anything that "wins" in that system much be more intelligent. Said differently, intelligence may lead to wealth but wealth doesn't imply intelligence.
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#477Earlier quoted context omitted.
but e=mc^2 is just an approximation e: nice, downvoted for knowing special relativity
Can you elaborate? How is E=mc^2 an approximation, in special relativity or otherwise? What is it an approximation of?
where p is momentum. When an object is traveling at relativistic speeds, the momentum forms a more significant portion of its energy
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#478Earlier quoted context omitted.
I've read the thread and I think it's not very coherent overall, I also not sure if we disagree =) I agree that having putnam problems on OpenAI training set is not a smoking gun, however it's (almost) certain they are on training set, and having them would affect performance of the model on them too. Hence research like this is important, since it shows that observed behavior of the models is memoization to large ex…
nobody serious (like OAI) was using the putnam problems to claim generalization. this is a refutation in search of a claim - and many people in the upstream thread are suggesting that OAI is doing something wrong by training on a benchmark. OAI uses datasets like frontiermath or arc-agi that are actually held out to evaluate generalization.
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#479Earlier quoted context omitted.
So do what the commenter suggests and make irrelevant permutations to the input to find when it fails. ie., engage in hypothesis testing rather than confirmation bias. If a system has the capability to solve problems of {parts1...parts_n}, then it only has that capability if irrelevant permutations {parts1..parts2'...parts_n} make no difference. Its very obvious that such permutations can destory such apparent capabi…
I've just tested a number of permutations with Claude 3.5 Sonnet. It correctly answered all variants I tried on the first attempt, as follows: Which is heavier, a 9.99 kilogram tungsten cube or a 10.01 kilogram block of aerogel? Which is heavier, 10,000 steel balls weighing 0.999 grams each or 10,000 polystyrene balls weighing 1.001 grams each? Which is heavier, a 10.01kg block of steel on Venus or a 9.99kg bag of fe…
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#480Earlier quoted context omitted.
That appears to be the same model I used. This is why I emphasized I didn't "go shopping" for a result. That was the first result I got. I'm not at all surprised that it will nondeterministically get it correct sometimes. But if it doesn't get it correct every time, it doesn't "know". (In fact "going shopping" for errors would still even be fair. It should be correct all the time if it "knows". But it would be differ…
I don't believe that is the model that you used. I wrote a script and pounded 01 mini and gpt 4 with a wide vareity of tempature and top_p parameters, and was unable to get it to give the wrong answer a single time. Just a whole bunch of: (openai-example-py3.12) :~/code/openAiAPI$ python3 featherOrSteel.py Response 1: A 10.01-pound bag of fluffy cotton is heavier than a 9.99-pound bag of steel ingots. Response 2: A 1…