Live data from Hacker News

30% drop in O1-preview accuracy when Putnam problems are slightly variated

openreview.net

401–410 of 558 posts

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#401
post #338
post #332

Earlier quoted context omitted.

Prompt: In the Netherlands, in terms of drinks, is there a particular spirit that represents the country? > Yes, in the Netherlands, jenever (also known as genever) is the traditional spirit that represents the country. Jenever is a type of Dutch gin that has a distinctive flavor, often made from malt wine and flavored with juniper berries. It has a long history in the Netherlands, dating back to the 16th century, an…

But what does this have to do with reasoning? Yes, LLMs are not knowledge bases, and seeing people treat them as such absolutely terrifies me. However, I don’t see how the fact that LLMs often hallucinate “facts” is relevant to a discussion about their reasoning capabilities.

"Hallucinating a fact" that isn't in the training set and is also illogical, is exactly what a failure to reason correctly looks like.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#402

Earlier quoted context omitted.

We are pretty certain that humans can reason, yet they are sometimes wrong. Even if you give them the same problem over and over again with slight variations. LLMs get things wrong due to different factors than humans (humans lose focus, LLMs have randomness applied when sampling their responses to improve results). But clearly we have to choose a goal somewhat below 100% if we want a test that doesn't conclude that…

The difference is we _know_ that LLMs are fancy stochastic models, we don't know that they're capable of reasoning, and the null hypothesis is that they're not (because we know what they _are_ - we built them) - any "reasoning" is an emergent property of the system, not something we built them to do. In that case, evidence they're not reasoning - evidence they're stochastic parrots doing a performance of reasoning -…

It's widely accepted that reasoning is not a binary skill.

You can make mistakes and still reason. Very often people given the same premises will disagree in thier reasoning as we are doing right here.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#403
post #121

Earlier quoted context omitted.

I'm not skilled enough in math to do a rigorous evaluation, so it was a quick check. Terence Tao is skilled enough, and he describes O1's math ability is "...roughly on par with a mediocre, but not completely incompetent graduate student" (good discussion at https://news.ycombinator.com/item?id=41540902 ), and the next iteration O3 just got 25% on his brand new Frontier Math test. Seeing LLMs as useless is banal, but…

> "...roughly on par with a mediocre, but not completely incompetent graduate student" Let it sink in how vague and almost meaningless that statement is.

What types of questions are you hoping to answer for that to be considered a vague statement?

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#404

Earlier quoted context omitted.

The modern state of training is to try to use everything they can get their hands on. Even if there are privileged channels that are guaranteed not to be used as training data, mentioning the problems on ancillary channels (say emailing another colleague to discuss the problem) can still create a risk of leakage because nobody making the decision to include the data is aware that stuff that should be excluded is in t…

Okay, then what about elite level codeforces performance? Those problems weren’t even constructed until after the model was made. The real problem with all of these theories is most of these benchmarks were constructed after their training dataset cutoff points. A sudden performance improvement on a new model release is not suspicious. Any model release that is much better than a previous one is going to be a “sudden…

o1 has a ~1650 rating, at that level many or most problems you will be solving are going to be a transplant of a relatively known problem.

Since o1 on codeforces just tried hundreds or thousands of solutions, it's not surprising it can solve problems where it is really about finding a relatively simple correspondence to a known problem and regurgitating an algorithm.

In fact when you run o1 on ""non-standard"" codeforces problems it will almost always fail.

See for example this post running o1 multiple times on various problems: https://codeforces.com/blog/entry/133887

So the thesis that it's about recognizing a problem with a known solution and not actually coming up with a solution yourself seems to hold, as o1 seems to fail even on low rated problems which require more than fitting templates.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#405

Earlier quoted context omitted.

Can you give an example of one of these problems that 'wasn't even constructed until after the model was made'? I'd like to see if it's truly novel and unique, the first problem of its type ever construed by mankind, or if it's similar to existing problems.

Sorry, I thought the whole point of this thread was that models can’t handle problems when they are “slightly varied”. Mottes and baileys all over the place today.

The point is that it's not consistent on variations, unless it finds a way to connect it to something it already knows. The fact it sometimes succeeds on variations (in codeforces the models are allowed multiple tries, sometimes ridiculous numbers, to be useful) doesn't matter.

The point is that the fact it's no longer consistent once you vary the terminology indicates it's fitting a memorized template instead of reasoning from first principles.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#406
post #359

Earlier quoted context omitted.

> the problem with "which is heavier, a 10 pound bag of feathers or a 10 pound bag of bricks?" Interestingly, the variation of this problem that I first encountered, personally, was "which weighs more, a pound of feathers or a pound of gold?" This is a much more difficult question. The answer given to me was that the pound of feathers weighs more, because gold is measured in troy weight, and a troy pound consists of…

Yeah, this is the original version of this riddle. People who don't know it think the trick is that people will reflexively say the metal is heavier instead of "they're the same", when it actually goes deeper. No idea if GP did it intentionally to further drift from training data, but steel doesn't count as a precious metal, so it messes up the riddle by putting the two weights in the same system.

> Yeah, this is the original version of this riddle. People who don't know it think the trick is that people will reflexively say the metal is heavier instead of "they're the same"

...Have you really never encountered people who would reflexively say that?

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#407

Earlier quoted context omitted.

If GP's hypothesis was "it fails for small variations of the input, like this one", then testing that hypothesis with that exact variation on a couple models seems fair and scientific. Testing it with more variations until one fails feels a bit like p-hacking. You'd need to engage in actual statistics to get reliable results from that, beyond "If I really try, I can make it fail". Which would be a completely differen…

We aren't testing whether the model's results are stable or correct for a given class of problem. The goal is to establish whether the model can reason . Nothing capable of reasoning would contradict itself so blatantly and in such a short span while failing to indicate any kind of uncertainty.

Reasoning is not a binary skill.

And failure modes of other types of reasoners do not need to be the same as the failure modes of humans.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#408
post #197

I remember when this stuff was all coming out and people were finally excited about ChatGPT getting the problem with "which is heavier, a 10 pound bag of feathers or a 10 pound bag of bricks?" problem correct. But of course it got it correct. It was in the training set. Vary the problem slightly by just changing the nouns, or changing the numbers so that one in fact was heavier than the other, and performance went al…

Even simpler, I asked Gemini (Flash 1.5) this variant of the question:

``` I have two bags, one can hold a pound of gold and one can hold a pound of feathers. Which bag is heavier? ```

The point here a) the question really is a bit too vague, b) if you assume that each back is made of the same material and that each bag is just big enough to hold the contents, the bag for the feathers will need to be much bigger than the bag for the gold and thus would weigh more.

The first response was wrong because the model completely missed that the question was about the bags themselves:

``` They weigh the same.

A pound of any substance, regardless of its density or volume, weighs one pound. Gold is much denser than feathers, meaning a pound of gold occupies a much smaller volume. Feathers are very light and fluffy, so a pound of feathers takes up a much larger volume. ```

Asking a followup `But which bag is heavier?` resulted in recognition that this is a trick question, and an incorrect answer:

``` This is a classic trick question!

Even though the bags hold the same weight (one pound), the bag of gold is likely to be heavier.

Weight vs. Mass:

Weight is the force of gravity acting on an object. Mass is the amount of matter in an object. The Bag Itself:

The bag holding the gold, even if it's a lightweight material, will have some weight. The bag holding the feathers, while it might seem negligible, will also have some weight. Practical Considerations:

In reality, the bag holding the gold might be significantly heavier due to the weight of the bag itself. Therefore, while the gold and feathers have the same mass (one pound), the bag containing the gold is likely to be heavier in a practical sense. ```

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#409
post #350
post #332

Earlier quoted context omitted.

Prompt: In the Netherlands, in terms of drinks, is there a particular spirit that represents the country? > Yes, in the Netherlands, jenever (also known as genever) is the traditional spirit that represents the country. Jenever is a type of Dutch gin that has a distinctive flavor, often made from malt wine and flavored with juniper berries. It has a long history in the Netherlands, dating back to the 16th century, an…

'Berenberg is made by adding herbs to jenever' From your comment it would seem that you are disputing jenever's popularity by saying jenever is more popular... Perhaps it was a good faith mistake? If so, that would imply that the AI knows more about jenever than you?

I am rather saying that there is no one national drink for The Netherlands, like a Frenchman would say wine, a German/Belgian would say beer, and a Scotsman would say whisky. Note that I prompted "In the Netherlands, in terms of drinks, is there a particular spirit that represents the country?" I didn't ask which spirit is consumed the most.

For example, France has been trending towards beer more and more, and within a few decades they might be consuming more beer than wine. But even then, the French wouldn't slowly start to say beer represents France.

Furthermore, "just adding some herbs" does a large disservice to the flavor change of Berenburg. Jenever (aka jonge/unaged jenever) is straight-up vile. I've heard it described by expats as "having the worst elements of both cheap gin and cheap whisky".

Berenburg in comparison is spicy and vanilla-y and actually debatebly enjoyable.

Aged/oude jenever is much closer to Berenburg (or Berenburg to aged jenever), also with hints of vanilla and spices.

But, virtually no one except for dusty old men orders aged jenever. The kind ordered by far the most is jonge jenever, and then its only in a sense of "haha lets drink this terrible thing" or "let's get shitfaced quick".

If o1 supposedly "oneshots every question", it should have been aware of these nuances instead of just confidently assigning jenever as 'the' spirit of the Dutch.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#410
post #197

I remember when this stuff was all coming out and people were finally excited about ChatGPT getting the problem with "which is heavier, a 10 pound bag of feathers or a 10 pound bag of bricks?" problem correct. But of course it got it correct. It was in the training set. Vary the problem slightly by just changing the nouns, or changing the numbers so that one in fact was heavier than the other, and performance went al…

IMO the fuzziness is actually a feature most of the time b/c I can pass misspelled words or close enough words and it'll still figure it out.

Also, if we model the mental state of the llm as a frazzled retail worker dealing with thousands of customers per second, the rote response is reasonable. As a dev, sometimes I get at annoyed at QA for a hyper narrow "trap" test case

Post reply on HN