Live data from Hacker News

30% drop in O1-preview accuracy when Putnam problems are slightly variated

openreview.net

331–340 of 558 posts

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#331

Earlier quoted context omitted.

No, it only says "ChatGPT Plus (Upgrade)" or "ChatGPT". Maybe it's different if you have a paid account?

if i go to chatgpt.com on my phone not logged on at all it tells me very prominently at the top that i am using 4o mini

Someone else in this thread said,

> With a free account the model it claims to be using is "4o auto", which is not a model but apparently an attempt to automatically decide models for you to be more cost effective.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#332

Earlier quoted context omitted.

https://chatgpt.com/share/67756897-8974-8010-a0e0-c9e3b3e91f... so far o1-mini has bodied every task people are saying LLMs can’t do in this thread

This happens literally every time. Someone always says "ChatGPT can't do this!", but then when someone actually runs the example, chatGPT gets it right. Now what the OP is going to do next is proceed to move goalposts and say like "but umm I just asked chatgpt this, so clearly they modified the code in realtime to get the answer right"

Prompt: In the Netherlands, in terms of drinks, is there a particular spirit that represents the country?

> Yes, in the Netherlands, jenever (also known as genever) is the traditional spirit that represents the country. Jenever is a type of Dutch gin that has a distinctive flavor, often made from malt wine and flavored with juniper berries. It has a long history in the Netherlands, dating back to the 16th century, and is considered the precursor to modern gin.

> Jenever comes in two main types: oude (old) jenever, which has a richer, maltier flavor and is often aged in wooden casks, and jonge (young) jenever, which is lighter and has a more neutral taste. Both types can be enjoyed straight or used in cocktails.

> In addition to jenever, the Netherlands is also known for its liqueurs, such as advocaat, a rich and creamy drink made from eggs, sugar, and brandy, often enjoyed as a dessert or in cocktails. However, jenever remains the most iconic spirit associated with Dutch culture.

This is completely wrong. Jenever certainly is very Dutch, but no one would say it is iconic as the Dutch spirit. For example, if you asked up north in Friesland, they would say Berenburg.

This happens literally every time. Someone always says "ChatGPT can do this!", but then within one or two prompts, its gets it wrong.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#333
post #197

I remember when this stuff was all coming out and people were finally excited about ChatGPT getting the problem with "which is heavier, a 10 pound bag of feathers or a 10 pound bag of bricks?" problem correct. But of course it got it correct. It was in the training set. Vary the problem slightly by just changing the nouns, or changing the numbers so that one in fact was heavier than the other, and performance went al…

[flagged]

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#334

Earlier quoted context omitted.

So do what the commenter suggests and make irrelevant permutations to the input to find when it fails. ie., engage in hypothesis testing rather than confirmation bias. If a system has the capability to solve problems of {parts1...parts_n}, then it only has that capability if irrelevant permutations {parts1..parts2'...parts_n} make no difference. Its very obvious that such permutations can destory such apparent capabi…

> ...engage in hypothesis testing rather than confirmation bias Please leave the premises, sir. We don't take kindly to luddites here.

Tough crowd

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#335
post #197

I remember when this stuff was all coming out and people were finally excited about ChatGPT getting the problem with "which is heavier, a 10 pound bag of feathers or a 10 pound bag of bricks?" problem correct. But of course it got it correct. It was in the training set. Vary the problem slightly by just changing the nouns, or changing the numbers so that one in fact was heavier than the other, and performance went al…

10 pounds of bricks is actually heavier than 10 pounds of feathers.

Can you explain?

An ounce of gold is heavier than an ounce of feathers, because the "ounce of gold" is a troy ounce, and the "ounce of feathers" is an avoirdupois ounce. But that shouldn't be true between feathers and bricks - they're both avoirdupois.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#336
post #323

Earlier quoted context omitted.

yeah it gets it https://chatgpt.com/share/67757720-3c7c-8010-a3e9-ce66fb9f17... e: cool, this gets downvoted

It got it right, but an interesting result that it rambled on about monetary value for... no reason. > While the lint bag is heavier in terms of weight, it's worth mentioning that gold is significantly more valuable per pound compared to lint. This means that even though the lint bag weighs more, the gold bag holds much greater monetary value.

Legal said someone might sell a bag of gold for one of lint without it.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#337
post #288

Earlier quoted context omitted.

> That addition is a small capability but you only need a single counterexample to disprove a theory No, that's not how this works :) You can hardcode an exception to pattern recognition for specific cases - it doesn't cease to be a pattern recognizer with exceptions being sprinkled in. The 'theory' here is that a pattern recognizer can lead to AGI. That is the theory. Someone saying 'show me proof or else I say a pa…

It's not hardcoded, reissbaker has addressed this point. I think you are misinterpreting what the argument is. The argument being made is that LLMs are mere 'stochastic parrots' and therefore cannot lead to AGI. The analogy to Russell's teapot is that someone is claiming that Russells teapot is not there because china cannot exist in the vacuum of space. You can disprove that with a single counterexample. That does n…

[deleted]

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#338
post #332

Earlier quoted context omitted.

This happens literally every time. Someone always says "ChatGPT can't do this!", but then when someone actually runs the example, chatGPT gets it right. Now what the OP is going to do next is proceed to move goalposts and say like "but umm I just asked chatgpt this, so clearly they modified the code in realtime to get the answer right"

Prompt: In the Netherlands, in terms of drinks, is there a particular spirit that represents the country? > Yes, in the Netherlands, jenever (also known as genever) is the traditional spirit that represents the country. Jenever is a type of Dutch gin that has a distinctive flavor, often made from malt wine and flavored with juniper berries. It has a long history in the Netherlands, dating back to the 16th century, an…

But what does this have to do with reasoning? Yes, LLMs are not knowledge bases, and seeing people treat them as such absolutely terrifies me. However, I don’t see how the fact that LLMs often hallucinate “facts” is relevant to a discussion about their reasoning capabilities.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#339
post #197

I remember when this stuff was all coming out and people were finally excited about ChatGPT getting the problem with "which is heavier, a 10 pound bag of feathers or a 10 pound bag of bricks?" problem correct. But of course it got it correct. It was in the training set. Vary the problem slightly by just changing the nouns, or changing the numbers so that one in fact was heavier than the other, and performance went al…

Add some extra information, and it gets confused. This is 4o.

https://chatgpt.com/share/67759723-f008-800e-b0f3-9c81e656d6...

One might argue that it's impossible to compress air using known engineering, but that would be a different kind of answer.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#340

Earlier quoted context omitted.

There doesn't seem to be a way to choose a model up-front with a free account, but after you make a query you can click on the "regenerate" button and select whether to try again with "auto", 4o, or 4o-mini. At least until you use 4o too many times and get rate limited.

you can select the model in the header bar when you start a chat: the name of the currently selected model can be clicked to reveal a dropdown

Are you on the free version? Because for me it did not show there, only on the paid one.
Post reply on HN