Live data from Hacker News

30% drop in O1-preview accuracy when Putnam problems are slightly variated

openreview.net

311–320 of 558 posts

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#311
post #197

I remember when this stuff was all coming out and people were finally excited about ChatGPT getting the problem with "which is heavier, a 10 pound bag of feathers or a 10 pound bag of bricks?" problem correct. But of course it got it correct. It was in the training set. Vary the problem slightly by just changing the nouns, or changing the numbers so that one in fact was heavier than the other, and performance went al…

> the problem with "which is heavier, a 10 pound bag of feathers or a 10 pound bag of bricks?"

Interestingly, the variation of this problem that I first encountered, personally, was "which weighs more, a pound of feathers or a pound of gold?"

This is a much more difficult question. The answer given to me was that the pound of feathers weighs more, because gold is measured in troy weight, and a troy pound consists of only 12 ounces compared to the 16 ounces in a pound avoirdupois.

And that's all true. Gold is measured in troy weight, feathers aren't, a troy pound consists of only 12 ounces, a pound avoirdupois consists of 16, and a pound avoirdupois weighs more than a troy pound does.

The problem with this answer is that it's not complete; it's just a coincidence that the ultimate result ("the feathers are heavier") is correct. Just as a pound avoirdupois weighs more than a troy pound, an ounce avoirdupois weighs less than a troy ounce. But this difference, even though it goes in the opposite direction, isn't enough to outweigh the difference between 16 vs 12 ounces per pound.

Without acknowledging the difference in the ounces, the official answer to the riddle is just as wrong as the naive answer is.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#312
post #204

Earlier quoted context omitted.

That appears to be the same model I used. This is why I emphasized I didn't "go shopping" for a result. That was the first result I got. I'm not at all surprised that it will nondeterministically get it correct sometimes. But if it doesn't get it correct every time, it doesn't "know". (In fact "going shopping" for errors would still even be fair. It should be correct all the time if it "knows". But it would be differ…

you sure? i just asked o1-mini ( not 4o mini ) 5 times in a row (new chats obviously) and it got it right every time perhaps you stumbled on a rarer case but reading the logs you posted this sounds more like a 4o model than an o1 because it’s doing its thinking in the chat itself plus the procedure you described would probably get you 4o-mini

May be unrelated, but I have been having a lot of issues lately with ChatGPT letting me select a model (o1) and silently switching to 4o.

This is coming off my TWO DAY cooldown on o1 usage, which is frustrating.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#313
post #197

I remember when this stuff was all coming out and people were finally excited about ChatGPT getting the problem with "which is heavier, a 10 pound bag of feathers or a 10 pound bag of bricks?" problem correct. But of course it got it correct. It was in the training set. Vary the problem slightly by just changing the nouns, or changing the numbers so that one in fact was heavier than the other, and performance went al…

Are you sure you weren't fishing? I ran 5 sessions and never got the wrong answer. All using gpt 4o-mini, which is the default non logged in experience on chatgpt.com.

1. The 10.01-pound bag of fluffy cotton is heavier than the 9.99-pound bag of steel ingots. Despite the difference in material density, the key factor here is the weight itself, with 10.01 pounds being greater than 9.99 pounds, regardless of the substance.

2. The 10.01-pound bag of fluffy cotton is heavier than the 9.99-pound bag of steel ingots. The weight of an object is directly determined by its mass, and in this case, the cotton bag weighs 10.01 pounds, while the steel ingot bag weighs 9.99 pounds, making the cotton bag heavier by 0.02 pounds.

3. The 10.01-pound bag of fluffy cotton is heavier than the 9.99-pound bag of steel ingots, even though the cotton is much less dense. The comparison is purely based on weight, and 10.01 pounds is greater than 9.99 pounds.

4. The 10.01-pound bag of fluffy cotton is heavier than the 9.99-pound bag of steel ingots. While steel is much denser than cotton, the weight is the determining factor here, and 10.01 pounds is slightly more than 9.99 pounds. So, the cotton bag weighs more.

5. The 10.01-pound bag of fluffy cotton is heavier. Despite the cotton being less dense and bulkier, the weight of the cotton is still 10.01 pounds, which is greater than the 9.99-pound bag of steel ingots.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#314
post #256

Earlier quoted context omitted.

Not an LLM but models could get pretty good at weather https://www.technologyreview.com/2024/12/04/1107892/google-d...

No, they don't, since the weather is chaotic. I mean, there are the theorems about how close you can get, and models are not better than theoretically possible.

Yeah, I wish more people understood that it is simply not possible to make precise long-term forecasts of chaotic systems. Whether it is weather, financial markets, etc.

It is not that we don't know yet because our models are inadequate, it's that it is unknowable.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#315
post #204

Earlier quoted context omitted.

That appears to be the same model I used. This is why I emphasized I didn't "go shopping" for a result. That was the first result I got. I'm not at all surprised that it will nondeterministically get it correct sometimes. But if it doesn't get it correct every time, it doesn't "know". (In fact "going shopping" for errors would still even be fair. It should be correct all the time if it "knows". But it would be differ…

I don't believe that is the model that you used. I wrote a script and pounded 01 mini and gpt 4 with a wide vareity of tempature and top_p parameters, and was unable to get it to give the wrong answer a single time. Just a whole bunch of: (openai-example-py3.12) :~/code/openAiAPI$ python3 featherOrSteel.py Response 1: A 10.01-pound bag of fluffy cotton is heavier than a 9.99-pound bag of steel ingots. Response 2: A 1…

Downvoted for… too conclusively proving OP wrong?

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#316

Earlier quoted context omitted.

Just because you can generalize the topic doesn't mean you can ignore the specific conversation and choose your hill to argue. Additionally, the conversation of this topic is about the model's ability to generalize and it's potential overfitting, which is arguably more important than parroting mathematics.

performance on a held-out set (like frontiermath) compared to putnam (which is not held out) is obviously relevant to a model's potential overfitting. i'm not going to keep replying, others can judge whether they think what i'm saying is "relevant at all."

Again, you set your own goal posts and failed to add any insights.

The topic here isn't "o-series sucks", it's addressing a found concern.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#317

An interesting example of this is: There are 6 “a”s in the sentence: “How many ‘a’ in this sentence?” https://chatgpt.com/share/677582a9-45fc-8003-8114-edd2e6efa2... Whereas the typical “strawberry” variant is now correct. There are 3 “r”s in the word “strawberry.” Clearly the lesson wasn’t learned, the model was just trained on people highlighting this failure case.

It also fails on things that aren't actual words

For example, the output for "how many x's are there in xaaax" is 3.

https://chatgpt.com/share/677591fe-aa58-800e-9e7a-81870387be...

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#318

Earlier quoted context omitted.

> just asked o1-mini (not 4o mini) 5 times in a row (new chats obviously) and it got it right every time Could you try playing with the exact numbers and/or substances?

give me a query and i’ll ask it, but also i don’t want to burn through all of my o1mini allocation and have to use the pay-as-you-go API.

>> so far o1-mini has bodied every task people are saying LLMs can’t do in this thread

> give me a query and i’ll ask it

Here's a query similar to one that I gave to Google Gemini (version unknown), which failed miserably:

---query---

Steeleye Span's version of the old broadsheet ballad "The Victory" begins the final verse with these lines:

Here's success unto the Victory / and crew of noble fame

and glory to the captain / bold Nelson was his name

What does the singer mean by these lines?

---end query---

Italicization is for the benefit of HN; I left that out of my prompt.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#319
post #197

I remember when this stuff was all coming out and people were finally excited about ChatGPT getting the problem with "which is heavier, a 10 pound bag of feathers or a 10 pound bag of bricks?" problem correct. But of course it got it correct. It was in the training set. Vary the problem slightly by just changing the nouns, or changing the numbers so that one in fact was heavier than the other, and performance went al…

10 pounds of bricks is actually heavier than 10 pounds of feathers.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#320
post #197

I remember when this stuff was all coming out and people were finally excited about ChatGPT getting the problem with "which is heavier, a 10 pound bag of feathers or a 10 pound bag of bricks?" problem correct. But of course it got it correct. It was in the training set. Vary the problem slightly by just changing the nouns, or changing the numbers so that one in fact was heavier than the other, and performance went al…

ChatGPT Plus user here. The following are all fresh sessions and first answers, no fishing. GPT 4: The 10.01-pound bag of fluffy cotton is heavier than the 9.99-pound bag of steel ingots. The type of material doesn’t affect the weight comparison; it’s purely a matter of which bag weighs more on the scale. GPT 4o: The 10.01-pound bag of fluffy cotton is heavier. Weight is independent of the material, so the bag of cot…

[deleted]
Post reply on HN