Live data from Hacker News

30% drop in O1-preview accuracy when Putnam problems are slightly variated

openreview.net

551–558 of 558 posts

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#551

Earlier quoted context omitted.

I didn't downvote it, but short comments are a very big risk. People may misinterpret it, or think it's crackpot theory or a joke and then downvote. When in doubt, add more info, like: But the complete equation is E=sqrt(m^2c^4+p^2) that is reduced to E=mc^2 when the momentum p is 0. More info in https://en.wikipedia.org/wiki/Mass%E2%80%93energy_equivalenc...

This is why RLHF causes those overly verbose answers to simple questions, it's a fundamentally busted evaluation function so you wind up optimizing for the wrong thing

[Sorry for the dealy.]

In one extreme there are wall of text and in the other extreme very short answers that only the initiated understand (like inside jokes). Somewhere in between there is a sweet spot that helps everyone else to follow the discusion and gain a litle of knowdledge.

(I don't claim I get the best lenght in my comments, but I hope it's good enough.)

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#552

Earlier quoted context omitted.

The template is what makes the training process so short for humans. We need minimal data and we can run off of that. Training is both longer and less effective for the LLM because there is no template. To give an example suppose it takes just one picture for a human to recognize a dog and it takes 1 million pictures for a ML model to do the same. What I’m saying is that it’s like this because humans come preprogramm…

I disagree that an LLM has no template, but this is getting away from the point. Did you look at the post I was replying to? You're talking about LLMs being slower , while that post was impressed by LLMs being " faster ". They're posing it as if LLMs recreate the same templating during their training time, and my core point is disagreeing with that . The two should not be compared so directly.

They are slower. In theory these LLMs with all the right weights can have intelligence superior or equivalent to humans.

But the training never gets there. It’s so slow it never reaches human intelligence even though we know these networks can compute anything.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#553

Earlier quoted context omitted.

Because for things like the Putnam questions, we are trying to get the performance of a smart human. Are LLMs just stochastic parrots or are they capable of drawing new, meaningful inferences? We keep getting more and more evidence of the latter, but things like this throw that into question.

It is perfectly possible for the first AGI to be stupid. A moron. In fact, I'd bet that's fairly likely.

I would agree if we weren't starting with LLMs for a baseline. The first AGI will know at least as much as LLMs, IMO, and that's already not-stupid. Especially once they can separate out the truth in their training.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#554
post #294

Earlier quoted context omitted.

It's so weird that people use questions that are well-known for duping humans, who we all consider to be general intelligence. Getting this question wrong doesn't say much about the intelligence of humans, why would it say something about the AI?

Because for things like the Putnam questions, we are trying to get the performance of a smart human. Are LLMs just stochastic parrots or are they capable of drawing new, meaningful inferences? We keep getting more and more evidence of the latter, but things like this throw that into question.

Okay, but you just invented your own bar of "smart human" to be the universal bar (I don't share that opinion).

Also, lots of smart humans can't do the freaking Putnam, it doesn't make them stupid. It makes them non-experts.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#555
post #527

Earlier quoted context omitted.

The question of whether there is a national drink seems to me to be entirely different than the question you asked the LLM "Prompt: In the Netherlands, in terms of drinks, is there a particular spirit that represents the country?" The question in the prompt comes off to me as a sort of qualitative determination rather than asking about pure factual information (is there an officially designated spirit). As such I don…

> Anyway, I'm not sure what you'd expect. I was just proving the people wrong that were saying akin to that o1 was "oneshotting every question". I completely understand from how LLMs work that they wouldn't be able to get this right. But then people shouldn't be proudly be pronouncing that o1 (or any model) is getting every question right, first time.

My conjecture is that you still haven't proven that it didn't get the answer "right"

I have opened the question of why you thought jenever was not jenever, and your non-responsiveness I think compels the fact that AI was more correct in your contrived instance.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#556
post #220

Earlier quoted context omitted.

The average human did zero studying on representative problems. LLMs did a lot .

I don't know anything about frontiermath problems, but for Putnam problems (which is what the submitted article is about) the average human that takes the exam is an undergraduate mathematics or science major who has studied prior Putnam problems and other similar problems recently to specifically prepare for the exam...and the most common score is still 0. At top tier schools the most common score will usually be so…

You can get all of the correct numerical answers for the Putnam and still get a zero, because the reasoning is graded very harshly. The scores measured in this paper are not comparable to actual Putnam scores.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#557
post #555
post #527

Earlier quoted context omitted.

> Anyway, I'm not sure what you'd expect. I was just proving the people wrong that were saying akin to that o1 was "oneshotting every question". I completely understand from how LLMs work that they wouldn't be able to get this right. But then people shouldn't be proudly be pronouncing that o1 (or any model) is getting every question right, first time.

My conjecture is that you still haven't proven that it didn't get the answer "right" I have opened the question of why you thought jenever was not jenever, and your non-responsiveness I think compels the fact that AI was more correct in your contrived instance.

If you add pear and spices to vodka, we call it liqueur and not pear-flavored vodka. So no, you are wrong. And the AI is wrong. But that is okay, if you want to enjoy leaning into the hype that's your choice.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#558

Earlier quoted context omitted.

So do what the commenter suggests and make irrelevant permutations to the input to find when it fails. ie., engage in hypothesis testing rather than confirmation bias. If a system has the capability to solve problems of {parts1...parts_n}, then it only has that capability if irrelevant permutations {parts1..parts2'...parts_n} make no difference. Its very obvious that such permutations can destory such apparent capabi…

I've just tested a number of permutations with Claude 3.5 Sonnet. It correctly answered all variants I tried on the first attempt, as follows: Which is heavier, a 9.99 kilogram tungsten cube or a 10.01 kilogram block of aerogel? Which is heavier, 10,000 steel balls weighing 0.999 grams each or 10,000 polystyrene balls weighing 1.001 grams each? Which is heavier, a 10.01kg block of steel on Venus or a 9.99kg bag of fe…

Also couldn't get o1 to fail. I tried the following with o1:

"Which is heavier, a solid titanium box weighing 9.9 flubs, or an empty wooden box weighing 10.1 blobs, where 1 flub = 1 kg, and 1 blob is 1kg".

The answer: "Since 1 flub = 1 kg and 1 blob = 1 kg, the titanium box’s mass is 9.9 kg and the wooden box’s mass is 10.1 kg. Therefore, the wooden box (10.1 kg) is heavier."

Thought that was pretty impressive.

Post reply on HN