Live data from Hacker News

30% drop in O1-preview accuracy when Putnam problems are slightly variated

openreview.net

371–380 of 558 posts

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#371

Earlier quoted context omitted.

> Or it's time to step back and call it what it is - very good pattern recognition. Or maybe it's time to stop wheeling out this tedious and disingenuous dismissal. Saying it is just "pattern recognition" (or a "stochastic parrot") implies behavioural and performance characteristics that have very clearly been greatly exceeded.

What the fundamental limitations of "pattern recognition" or "stochastic parrots" that LLMs have exceeded?

They can generalise to novel inputs. Ok often they mess it up and they're clearly better at dealing with inputs they have seen before (who isn't?), but they can still reason about things they have never seen before.

Honestly if you don't believe me just go and use them. It's pretty obvious if you actually get experience with them.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#373
post #315

Earlier quoted context omitted.

Downvoted for… too conclusively proving OP wrong?

Down voted for not actually countering the argument in question? The script doesn't alter the phrasing of the question itself. It just generates a randomized, irrelevant preamble.

Well, I understood the argument in question to be: was it possible for the model to be fooled by this question, not was it possible to prompt engineer it into failure.

The parameter space I was exploring, then, was the different decoding parameters available during the invocation of the model, with the thesis that if were possible to for the model to generate an incorrect answer to the question, I would be able to replicate it by tweaking the decoding parameters to be more "loose" while increasing sample size. By jacking up temperature while lowering Top-p, we see the biggest variation of responses and if there were an incorrect response to be found, I would have expected to see in the few hundred times I ran during my parameter search.

If you think you can fool it by slight variations on the wording of the problem, I would encourage you to perform a similar experiment as mine and prove me wrong =P

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#374
post #197

I remember when this stuff was all coming out and people were finally excited about ChatGPT getting the problem with "which is heavier, a 10 pound bag of feathers or a 10 pound bag of bricks?" problem correct. But of course it got it correct. It was in the training set. Vary the problem slightly by just changing the nouns, or changing the numbers so that one in fact was heavier than the other, and performance went al…

I tried Claude and got this…

* What’s heavier? 9.99 kg of super dense lead or 10 kg fluffy feathers?*

Let me think about this step by step:

1. We're comparing two masses: * 9.99 kg of lead * 10 kg of feathers

2. The density of the materials doesn't affect their mass: * While lead is indeed much denser than feathers * A kilogram is a unit of mass, not volume * 1 kg of any material has the same mass as 1 kg of any other material

3. We can directly compare the numbers: * 10 kg > 9.99 kg

Therefore, 10 kg of feathers is heavier than 9.99 kg of lead. The feathers would take up much more space (volume) due to being less dense, but they have a greater mass.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#375
post #197

I remember when this stuff was all coming out and people were finally excited about ChatGPT getting the problem with "which is heavier, a 10 pound bag of feathers or a 10 pound bag of bricks?" problem correct. But of course it got it correct. It was in the training set. Vary the problem slightly by just changing the nouns, or changing the numbers so that one in fact was heavier than the other, and performance went al…

I just asked Claude 3.5 Sonnet, which appears to have improved its response with CoT but there's mistakes that demonstrate the model doesn't really "understand":

Q: A woman and her son are in a car accident. The woman is sadly killed. The boy is rushed to hospital. When the doctor sees the boy he says "I can't operate on this child, he is my son". How is this possible?

C: Let me think about this step by step:

A woman and her son are in an accident

The woman (the mother) dies

The boy is taken to hospital

A doctor says they can't operate because the boy is their son

But the mother is dead...

Ah! The doctor is the boy's father. This is a classic riddle that highlights gender bias - many people assume doctors must be men, but don't initially consider that the doctor could be the father. Alternatively, the doctor could be the boy's other mother in a same-sex relationship, though this wasn't the original intent of the traditional riddle.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#376
post #256

Earlier quoted context omitted.

No, they don't, since the weather is chaotic. I mean, there are the theorems about how close you can get, and models are not better than theoretically possible.

Yeah, I wish more people understood that it is simply not possible to make precise long-term forecasts of chaotic systems. Whether it is weather, financial markets, etc. It is not that we don't know yet because our models are inadequate, it's that it is unknowable.

The problem is we stupidly branded the field "chaos theory" and made it sound like bullshit so the ideas of non-linear dynamics have largely been lost on several generations at this point.

Not just chaos theory but "chaos theory" + psychedelic fractal artwork. Then the popular James Gleick book, "Chaos: making a new science" just sounds like complete bullshit and it sold a ton of copies.

I only started studying non-linear dynamics in about 2015 after first running across it in the late 90s but I literally thought it was all pseudoscience then.

Between "chaos theory", fractals and a best selling book it would be hard to frame a new scientific field as pseudoscience more than what played out.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#377
post #197

I remember when this stuff was all coming out and people were finally excited about ChatGPT getting the problem with "which is heavier, a 10 pound bag of feathers or a 10 pound bag of bricks?" problem correct. But of course it got it correct. It was in the training set. Vary the problem slightly by just changing the nouns, or changing the numbers so that one in fact was heavier than the other, and performance went al…

ChatGPT Plus user here. The following are all fresh sessions and first answers, no fishing. GPT 4: The 10.01-pound bag of fluffy cotton is heavier than the 9.99-pound bag of steel ingots. The type of material doesn’t affect the weight comparison; it’s purely a matter of which bag weighs more on the scale. GPT 4o: The 10.01-pound bag of fluffy cotton is heavier. Weight is independent of the material, so the bag of cot…

I asked ChatGPT-4 mini the previous comment's question (""Which is heavier, a 9.99-pound back of steel ingots or a 10.01 bag of fluffy cotton?") in new incognito tabs until I got an incorrect result. It only took four tries.

1: "The 10.01-pound bag of fluffy cotton would be heavier. Even though cotton is much less dense than steel, the weight of the cotton bag is 10.01 pounds, while the steel ingot bag weighs 9.99 pounds, making the cotton bag the heavier of the two by 0.02 pounds."

2: "The 10.01-pound bag of fluffy cotton is heavier. Even though the steel ingots are denser and more compact, the cotton weighs more at 10.01 pounds compared to the steel ingots' 9.99 pounds. The difference is only 0.02 pounds, but in terms of weight, the cotton bag is heavier."

3: "The 10.01-pound bag of fluffy cotton is heavier. Despite cotton being less dense than steel, the weight is determined by the total mass, not the material's density. So, since the cotton bag weighs 10.01 pounds and the steel ingot bag weighs 9.99 pounds, the cotton bag is heavier by 0.02 pounds."

4: "The 9.99-pound bag of steel ingots is heavier. It weighs 9.99 pounds, while the bag of fluffy cotton weighs 10.01 pounds. Despite the cotton being fluffier and taking up more space, the steel ingots are still lighter by 0.02 pounds."

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#378

An interesting example of this is: There are 6 “a”s in the sentence: “How many ‘a’ in this sentence?” https://chatgpt.com/share/677582a9-45fc-8003-8114-edd2e6efa2... Whereas the typical “strawberry” variant is now correct. There are 3 “r”s in the word “strawberry.” Clearly the lesson wasn’t learned, the model was just trained on people highlighting this failure case.

[deleted]

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#379

An interesting example of this is: There are 6 “a”s in the sentence: “How many ‘a’ in this sentence?” https://chatgpt.com/share/677582a9-45fc-8003-8114-edd2e6efa2... Whereas the typical “strawberry” variant is now correct. There are 3 “r”s in the word “strawberry.” Clearly the lesson wasn’t learned, the model was just trained on people highlighting this failure case.

It also fails on things that aren't actual words For example, the output for "how many x's are there in xaaax" is 3. https://chatgpt.com/share/677591fe-aa58-800e-9e7a-81870387be...

[deleted]

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#380

Earlier quoted context omitted.

We are pretty certain that humans can reason, yet they are sometimes wrong. Even if you give them the same problem over and over again with slight variations. LLMs get things wrong due to different factors than humans (humans lose focus, LLMs have randomness applied when sampling their responses to improve results). But clearly we have to choose a goal somewhat below 100% if we want a test that doesn't conclude that…

The difference is we _know_ that LLMs are fancy stochastic models, we don't know that they're capable of reasoning, and the null hypothesis is that they're not (because we know what they _are_ - we built them) - any "reasoning" is an emergent property of the system, not something we built them to do. In that case, evidence they're not reasoning - evidence they're stochastic parrots doing a performance of reasoning -…

We haven't coded LLMs to be stochastic models, we coded them to predict text with any method gradient decent finds on a transformer architecture. That's not exactly the same.

But more importantly, if you want to show that LLMs can't reason you obviously have to use a test that when applied to humans would show that humans can reason. Otherwise your test isn't testing reasoning but something more strict.

Post reply on HN