Live data from Hacker News

30% drop in O1-preview accuracy when Putnam problems are slightly variated

openreview.net

291–300 of 558 posts

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#291

Earlier quoted context omitted.

> just asked o1-mini (not 4o mini) 5 times in a row (new chats obviously) and it got it right every time Could you try playing with the exact numbers and/or substances?

give me a query and i’ll ask it, but also i don’t want to burn through all of my o1mini allocation and have to use the pay-as-you-go API.

> What is heavier a liter of bricks or a liter of feathers?

>> A liter of bricks and a liter of feathers both weigh the same—1 kilogram—since they each have a volume of 1 liter. However, bricks are much denser than feathers, so the bricks will take up much less space compared to the large volume of feathers needed to make up 1 liter. The difference is in how compactly the materials are packed, but in terms of weight, they are identical.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#292

Earlier quoted context omitted.

give me a query and i’ll ask it, but also i don’t want to burn through all of my o1mini allocation and have to use the pay-as-you-go API.

> What is heavier a liter of bricks or a liter of feathers? >> A liter of bricks and a liter of feathers both weigh the same—1 kilogram—since they each have a volume of 1 liter. However, bricks are much denser than feathers, so the bricks will take up much less space compared to the large volume of feathers needed to make up 1 liter. The difference is in how compactly the materials are packed, but in terms of weight,…

https://chatgpt.com/share/677583a3-526c-8010-b9f9-9b2a3374da... o1-mini best-of-1

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#293

Earlier quoted context omitted.

FYI: If you do that without a subscrpition, you currently (most likely) get a response generated through 4o-mini — which is not any of their reasoning models (o1, o1-mini or previously o1-preview) of the branch discussed in the linked paper. Notably, it's not even necessarily 4o, their premiere "non-reasoning"-model, but likely the cheaper variant: With a free account the model it claims to be using is "4o auto", whi…

There doesn't seem to be a way to choose a model up-front with a free account, but after you make a query you can click on the "regenerate" button and select whether to try again with "auto", 4o, or 4o-mini. At least until you use 4o too many times and get rate limited.

Ah, interesting!

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#294
post #204

Earlier quoted context omitted.

https://chatgpt.com/share/67756897-8974-8010-a0e0-c9e3b3e91f... so far o1-mini has bodied every task people are saying LLMs can’t do in this thread

That appears to be the same model I used. This is why I emphasized I didn't "go shopping" for a result. That was the first result I got. I'm not at all surprised that it will nondeterministically get it correct sometimes. But if it doesn't get it correct every time, it doesn't "know". (In fact "going shopping" for errors would still even be fair. It should be correct all the time if it "knows". But it would be differ…

It's so weird that people use questions that are well-known for duping humans, who we all consider to be general intelligence.

Getting this question wrong doesn't say much about the intelligence of humans, why would it say something about the AI?

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#295

Earlier quoted context omitted.

A machine that synthesizes all human knowledge really ought to know more than an individual in terms of intellect. An entity with all of human intellect prior to 1905 does not need to be as intelligent as a human to make discoveries that mere humans with limited intellect made. Why lower the bar?

"Why lower the bar?" Because of the chance of misundertanding. Failing at acknowledging artificial general intelligence standing right next to us. An incredible risk to take in alignment. Perfect memory doesn't equal to perfect knowledge, nor perfect understanding of everything you can know. In fact, a human can be "intelligent" with some of his own memories and/or knowledge, and - more commmonly - a complete "fool"…

> If we'd have TODAY an AI with the same parameters as the human with perfect memory, it will be most certainly closely examined and determined to be not a general artificial intelligence.

The human could learn to master a task, current AI can't. That is very different, the AI doesn't learn to remember stuff they are stateless.

When I can take an AI and get it to do any job on its own without any intervention after some training then that is AGI. The person you mentioned would pass that easily. Current day AI aren't even close.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#296

Earlier quoted context omitted.

it’s extremely easy to see which model you are using. one’s own… difficulties understanding are not a conspiracy by OpenAI

It does not show the model version anywhere on the page on chatgpt.com, even when logged in.

Yes it does, at the top of every chat there is a drop-down to select the model, which displays the current model. It's been a constant part of the UI since forever.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#297

ok but preview sucks, run it on o1 pro. 99% of studies claiming some out of distribution failure of an LLM uses a model already made irrelevant by SOTA. These kinds of studies, with long throughputs and review periods, are not the best format to make salient points given the speed at which the SOTA horizon progresses

I wonder what is baseline OOD generalization for humans. It takes around 7 years to generalize visual processing to X-ray images. How well does a number theorist respond to algebraic topology questions? How long it will take a human to learn to solve ARC challenges in the json format just as well as in the visual form?

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#298
post #47

Earlier quoted context omitted.

That is the principle behind the game 'Simon says'

That’s a very silly analogy. A more realistic analogy would be do humans perform better on computing 37x41 or 87x91 (with showing the work)?

It was not an analogy at all. It was an simplified example of the idea that a slight change in a pattern can induce error in humans.

It seems some people disagree that that is what the game "Simon Says" is about. I feel like they might play a vastly simplified version of the game that I am familiar with.

There was a recent episode of Game Changer based on this which is an excellent example of how the game leader should attempt to induce errors by making a change that does not get correctly accounted for.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#299
post #294
post #204

Earlier quoted context omitted.

That appears to be the same model I used. This is why I emphasized I didn't "go shopping" for a result. That was the first result I got. I'm not at all surprised that it will nondeterministically get it correct sometimes. But if it doesn't get it correct every time, it doesn't "know". (In fact "going shopping" for errors would still even be fair. It should be correct all the time if it "knows". But it would be differ…

It's so weird that people use questions that are well-known for duping humans, who we all consider to be general intelligence. Getting this question wrong doesn't say much about the intelligence of humans, why would it say something about the AI?

Because for things like the Putnam questions, we are trying to get the performance of a smart human. Are LLMs just stochastic parrots or are they capable of drawing new, meaningful inferences? We keep getting more and more evidence of the latter, but things like this throw that into question.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#300
post #197

I remember when this stuff was all coming out and people were finally excited about ChatGPT getting the problem with "which is heavier, a 10 pound bag of feathers or a 10 pound bag of bricks?" problem correct. But of course it got it correct. It was in the training set. Vary the problem slightly by just changing the nouns, or changing the numbers so that one in fact was heavier than the other, and performance went al…

As long as an LLM is capable of inserting "9.99 > 10.01?" into an evaluation tool, we're on a good way.

It feels a bit like "if all you have is a hammer, everything looks like a nail", where we're trying to make LLMs do stuff which it isn't really designed to do.

Why don't we just limit LLMs to be an interface to use other tools (in a much more human way) and train them to be excellent at using tools. It would also make them more energy efficient.

But it's OK if we currently try to make them do as much as possible, not only to check where the limits are, but also to gain experience in developing them and for other reasons. We just shouldn't expect them to be really intelligent.

Post reply on HN