Live data from Hacker News

30% drop in O1-preview accuracy when Putnam problems are slightly variated

openreview.net

451–460 of 558 posts

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#451

Earlier quoted context omitted.

> Good students are immune to variations I don't believe that. I'd put some good money that if an excellent student is given an exact question from a previous year, they'll do better (faster & more accurate) on it, than when they're given a variation of it.

The point is that the “good student” will still do well on the variations, not suffer a 30% decrease in grade.

I'm not following, why would you assume that a good student taking an exam at the edge of their ability not do significantly better if they trained on the exact same questions (with the same solutions), as opposed to ones that are slightly different? I for one have absolutely struggled as a student when faced with questions that seemed similar to ones from previous exams, but actually had a crucial difference, and by looking at the examples on pages 9&10 in the article, I'm pretty sure I would have been likely to be confused too.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#452
post #429

Earlier quoted context omitted.

"Hallucinating a fact" that isn't in the training set and is also illogical, is exactly what a failure to reason correctly looks like.

Reasoning involves making accurate inferences based on the information provided in the current context, rather than recalling arbitrary facts from the training data.

Yes, that's what I said. The whole point of hallucinations is that they aren't "arbitrary facts recalled from the training data". They represent attempts to synthesize (i.e., infer) new facts. But because the inferences are not accurate, and because the synthesis process is not sound, the attempt cannot be called reasoning.

It is equally possible to "reason" about things you already know, as about things you've just been told. In fact, the capacity to speculatively, without prompting attempt such reasoning is a big part of cognition.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#453
post #420

Earlier quoted context omitted.

I must be to tired as I can't find any flaw in that reasoning.

The joke/riddle text is "he says" but Claude says "their son" and suggests the doctor might be a woman. More substantively: "This is a classic riddle that highlights gender bias - many people assume doctors must be men, but don't initially consider that the doctor could be the father." is totally nonsensical. The text is a gender (and meaning) inversion of the classic riddle to confuse LLMs. Even though Claude correc…

Except that Claude often takes into account things it thinks might be typos.

This is not code. Forgetting a semi colon will not make the output break. It thinks 'maybe they wrote he instead of she' and then gives options for both situations.

It is meant to solve real world situations where people might not type properly, it is not a word problem solving machine.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#454
post #160

Earlier quoted context omitted.

Right. If it were trained through August 1939, how much prompting would be necessary to get it to predict aspects of WWII.

But we know Hitler has a Time Machine that goes forward, he doesn’t need to return to use that knowledge as he already has a timeline here to use. Definitely risks involved here.

If you build an oracle that tells you who wins the war that far in the future, you build a simulator that allows anyone to win any war. Everything is dual use.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#455
post #359

Earlier quoted context omitted.

Yeah, this is the original version of this riddle. People who don't know it think the trick is that people will reflexively say the metal is heavier instead of "they're the same", when it actually goes deeper. No idea if GP did it intentionally to further drift from training data, but steel doesn't count as a precious metal, so it messes up the riddle by putting the two weights in the same system.

> Yeah, this is the original version of this riddle. People who don't know it think the trick is that people will reflexively say the metal is heavier instead of "they're the same" ...Have you really never encountered people who would reflexively say that?

That's not what I said. I'm talking about the riddle itself, not how people react to it.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#456
post #368

Earlier quoted context omitted.

I’m not really sure what you’re trying to say here - that LLMs don’t work like human brains? We don’t need to conduct any analyses to know that LLMs don’t “know” anything in the way humans “know” things because we know how LLMs work. That doesn’t mean that LLMs aren’t incredibly powerful; it may not even mean that they aren’t a route to AGI.

>We don’t need to conduct any analyses to know that LLMs don’t “know” anything in the way humans “know” things because we know how LLMs work. People, including around HN, constantly argue (or at least phrase their arguments) as if they believed that LLMs do, in fact, possess such "knowledge". This very comment chain exists because people are trying to defend against a trivial example refuting the point - as if there…

You don't have to listen to or engage with those people though, just ignore 'em. People say all kinds of things on the Internet. It's completely futile to try to argue with or "correct" them all.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#457

Earlier quoted context omitted.

Could you share the exact chat you used for when it failed? There is a share chat button on openai. It's very difficult to be an AI bull when the goalposts are moving so quickly that ai answering core correctly across multiple models is brushed off as 'nondeterministically getting it correct sometimes'

Why? Did a grocery store self checkout ever fail to calculate sales tax? Do I need to run a study on that? The people selling this could not make a car drive but now its AGI.

A single-purpose state machine not failing to do the single thing it was created to do does not make for the clever retort you think it makes.

"AGI": emphasis on "G" for "General". The LLMs are not failing to do generalized tasks, and that they are nondeterministic is not a bug. Just don't use them for calculating sales tax. You wouldn't hire a human to calculate sales tax in their head, so why do you make this a requirement in order to call an LLM "AGI"?

I wonder when the goalposts will stop moving from "We have superhuman intelligences which are able to rather reliably converse in many languages, do generalized tasks and automate operations we thought were impossible to automate 3 years ago" (and by the way, this is what we have TODAY), all the way to "It's not AGI unless it's an omnipotent god that knows how to turn water into wine and calculate the applicable sales tax of that operation".

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#458
post #409
post #350

Earlier quoted context omitted.

'Berenberg is made by adding herbs to jenever' From your comment it would seem that you are disputing jenever's popularity by saying jenever is more popular... Perhaps it was a good faith mistake? If so, that would imply that the AI knows more about jenever than you?

I am rather saying that there is no one national drink for The Netherlands, like a Frenchman would say wine, a German/Belgian would say beer, and a Scotsman would say whisky. Note that I prompted "In the Netherlands, in terms of drinks, is there a particular spirit that represents the country?" I didn't ask which spirit is consumed the most. For example, France has been trending towards beer more and more, and within…

The question of whether there is a national drink seems to me to be entirely different than the question you asked the LLM "Prompt: In the Netherlands, in terms of drinks, is there a particular spirit that represents the country?"

The question in the prompt comes off to me as a sort of qualitative determination rather than asking about pure factual information (is there an officially designated spirit). As such I don't think it can necessarily be right or wrong.

Anyway, I'm not sure what you'd expect. In terms of acquisition of knowledge, LLMs fundamentally rely on a written corpus. Their knowledge of information that is passed through casual spoken conversation is limited. Sure, as human beings, we rely a great deal on the latter. But for an LLM to lack access to that information means that it's going to miss out on cultural nuances that are not widely expressed in writing. Much in the same way that a human adult can live in a foreign country for decades, speaking their adopted language quite fluently, but if they don't have kids of their own, they might be quite ignorant of that country's nursery rhymes and children's games, simply because they were never part of their acquired vocabulary and experience.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#459

Earlier quoted context omitted.

o1 has a ~1650 rating, at that level many or most problems you will be solving are going to be a transplant of a relatively known problem. Since o1 on codeforces just tried hundreds or thousands of solutions, it's not surprising it can solve problems where it is really about finding a relatively simple correspondence to a known problem and regurgitating an algorithm. In fact when you run o1 on ""non-standard"" codefo…

o3 is what i’m referring to and it is 2700

It's extremely unlikely for o3 to have hit 2700 on live contests as such a rapid increase in score would have been noticed by the community. I can't find anything online detailing how contamination was avoided since it clearly wasn't run live, including in their video, and neither could I find details about the methodology (number of submissions being the big one, in contests you can also get 'hacked' esp. at a high level), problem selection, etc...

Additionally, people weren't able to replicate o1-mini results in live contests straightforwardly - often getting scores between 700 and 1200, which raises questions as for the methodology.

Perhaps o3 really is that good, but I just don't see how you can claim what you claimed for o3, we have no idea that the problems have never been seen, and the fact people find much lower Elo scores with o1/o1-mini with proper methodology raises even more questions, let alone conclusively proving these are truly novel tasks it's never seen.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#460
post #232

Earlier quoted context omitted.

The heightening of the bar is an attempt to deny that milestones were surpassed and to claim that LLMs are not intelligent. We had a threshold for intelligence. An LLM blew past it and people refuse to believe that we passed a critical milestone in creating AI. Everyone still thinks all an LLM does is regurgitate things. But a technical threshold for intelligence cannot have any leeway for what people want to believe…

Disagree. The AI we have is very useful for specific things. The pushback you see is not so much denying the milestones that have been surpassed, but rather the milestones that enthusiasts claim are near. And for good reason! Every time and in every field we’ve extrapolated an exponential-looking curve ad infinitum, it’s turned out to be S-shaped, and life goes on. > We had a threshold for intelligence. We’ve had man…

> It’s somewhat true we’re moving the goalposts. But the reason is not stubbornness, but rather that we can’t properly define and subcategorize what reason and intelligence really is.

Disagree. Intelligence is a word created by humans. The entire concept is made up and defined by humans. It is not some concept that exists outside of that. It is simply a collection of qualities and features we choose to define as a word “intelligent”. The universe doesn’t really have a category or a group of features that is labeled intelligent. Does it use logic? Does it have feelings? Can it talk? Can it communicate? We define the features and we choose to put each and every feature under a category called “intelligence”.

Therefore when we define the “Turing test” as a benchmark for intelligence and we then invalidate it, it is indeed stubbornness and a conscious choice to change a definition of a word we Originally made up in the first place.

What you don’t realize is this entire thing is a vocabulary problem. When we argue what is conscious or what is intelligent we are simply arguing for what features belong in what categories we made up. When the category has blurry or controversial boundaries it’s because we chose the definition to be fuzzy. These are not profound discussions. They are debates about language choice. We are talking About personal definitions and generally accepted definitions both of which are completely chosen and made up by us. It is not profound to talk about things that are simply arbitrary choices picked by humans.

That being said we are indeed changing the goal posts. We are evolving our own chosen definitions and we very well may eventually change the definition of intelligence to never include any form of thinking machine that is artificially created. The reason why we do this is a choice. We are saying, “hey these LLMs are not anything amazing or anything profound. They are not intelligent and I choose to believe this by changing and evolving my own benchmark for what is intelligent.”

Of course this all happens subconsciously based off of deeply rooted instincts and feelings. It’s so deep that it’s really hard to differentiate the instincts between rational thinking. When you think logically, “intelligence” is just a word with an arbitrary definition. An arbitrary category. But the instincts are so strong that you literally spent your entire life thinking that intelligence like god or some other common myth made up by humans is some concept that exists outside of what we make up. It’s human to have these instincts, that’s where religion comes from. What you don’t realize is that it’s those same instincts fueling your definition of what is “intelligent”.

Religious people move the goal posts too. When science establishes things in reality like the helio centricity of the solar system religious people need to evolve their beliefs in order to stay inline with reality. They often do this by reinterpreting the Bible. It’s deeply rooted instincts that prevent us from thinking rationally and it effects the great debate we are having now on “what is intelligence?”.

Post reply on HN