Live data from Hacker News

30% drop in O1-preview accuracy when Putnam problems are slightly variated

openreview.net

421–430 of 558 posts

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#421

An interesting example of this is: There are 6 “a”s in the sentence: “How many ‘a’ in this sentence?” https://chatgpt.com/share/677582a9-45fc-8003-8114-edd2e6efa2... Whereas the typical “strawberry” variant is now correct. There are 3 “r”s in the word “strawberry.” Clearly the lesson wasn’t learned, the model was just trained on people highlighting this failure case.

It also fails on things that aren't actual words For example, the output for "how many x's are there in xaaax" is 3. https://chatgpt.com/share/677591fe-aa58-800e-9e7a-81870387be...

Transformers are very bad at counting in one feed forward pass, you need to explicitly tell them to use a counter in autoregressive fashion like here:

https://chatgpt.com/share/6775cb37-4198-8007-82cb-e897220827...

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#422

Earlier quoted context omitted.

do you understand the difference between test data and train data? just reread this thread of comments

I don't know why I and you are getting downvoted. Sometimes, HN crowd is just unhinged against AI.

These models are trained in two steps: training base model and then uptraining it. First step includes as much data as possible, everything company can find. For Llama models it's 15T tokens, which is ~40 TB of data. No-one really puts an effort on splitting this data into train/test/eval (and it's not very achievable either). It's just as much data as possible.

So it's like 99.9999999% wrong to assume something public isn't on the train set, such as Putnam problems in this case. This is about it.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#423
post #163

Earlier quoted context omitted.

https://s.h4x.club/bLuNed45 - it's more crazy to me that my wife CAN in fact read this stuff easily, vs the fact that an LLM can't. (for anyone who doesn't feel like downloading the zip, here is a single image from the zip: https://s.h4x.club/nOu485qx )

“ From what the text shows, Henry Jenkins and his wife Caroline (the boy’s mother) are asking the Orphans Court to void an apprenticeship arrangement involving her minor son, James Timmons. They claim James—about 15 years old—was bound out as an apprentice without proper authority or the mother’s consent, and they cite Maryland law (an act from 1793 and its supplements) which they believe was not followed. They reque…

She said: "The idea is close the the details are wrong".

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#424

Earlier quoted context omitted.

A machine that synthesizes all human knowledge really ought to know more than an individual in terms of intellect. An entity with all of human intellect prior to 1905 does not need to be as intelligent as a human to make discoveries that mere humans with limited intellect made. Why lower the bar?

The heightening of the bar is an attempt to deny that milestones were surpassed and to claim that LLMs are not intelligent. We had a threshold for intelligence. An LLM blew past it and people refuse to believe that we passed a critical milestone in creating AI. Everyone still thinks all an LLM does is regurgitate things. But a technical threshold for intelligence cannot have any leeway for what people want to believe…

> The heightening of the bar is an attempt to deny that milestones were surpassed and to claim that LLMs are not intelligent.

That was never "the bar"; nobody denies that milestones have been surpassed; none of those milestones are relevant to the question of intelligence.

> We had a threshold for intelligence. An LLM blew past it and people refuse to believe

Have you ever actually looked at contemporary (to Turing) examples of what people thought "passing a Turing test" might look like? It's abundantly clear to me that we were simply wrong about what the output would have to look like in order to convince human judges in the 2020s.

Even examples from much more recently (see e.g. on http://www-logic.stanford.edu/seminar/1213/Hawke_TuringTest....) suggest a very different approach to the test than prompting ChatGPT and marveling at the technical accuracy of its prose.

(Exercise: ask an LLM to write a refutation to your comment from the perspective of a human AI skeptic. Notice the ways in which it differs from mine.)

> Everyone still thinks all an LLM does is regurgitate things.

No; people still think LLMs aren't intelligent. Because they aren't, and they cannot become so in principle. They can do many things that are clearly beyond "regurgitation" (as we would otherwise apply the word to computer programs), but none of those things are the result of intelligence. Producing a result that could plausibly come from an intelligent system does not, in fact, demonstrate that the actual system producing it is also intelligent. The https://en.wikipedia.org/wiki/Antikythera_mechanism wasn't intelligent, either, and applying a power source to turn the gears wouldn't have made it so, either.

> They don’t want to define an LLM as intelligent even if it meets the Turing test technical definition of intelligence so they change the technical definition.

The Turing Test was never a "technical definition" of intelligence. Turing's original paper (https://en.wikipedia.org/wiki/Computing_Machinery_and_Intell...) spoke of "thinking" rather than "intelligence". Besides, the "Imitation Game" is presented as a substitute problem exactly because "think" cannot be clearly enough defined for the purposes. The entire point:

> As Stevan Harnad notes,[7] the question has become "Can machines do what we (as thinking entities) can do?" In other words, Turing is no longer asking whether a machine can "think"; he is asking whether a machine can act indistinguishably[8] from the way a thinker acts. This question avoids the difficult philosophical problem of pre-defining the verb "to think" and focuses instead on the performance capacities that being able to think makes possible, and how a causal system can generate them.

But the usual processes of pop science seem to have created a folk wisdom that being able to pass a Turing test logically ought to imply intelligence. This idea is what has been disproven, not the AI skepticism.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#425
post #163

Earlier quoted context omitted.

my wish for new years is that every time people make a comment like this they would share an example task

https://s.h4x.club/bLuNed45 - it's more crazy to me that my wife CAN in fact read this stuff easily, vs the fact that an LLM can't. (for anyone who doesn't feel like downloading the zip, here is a single image from the zip: https://s.h4x.club/nOu485qx )

One stumbling block is that OpenAI's best model (o1-pro) can't yet deal with image uploads (edit: I'm wrong, it works now... trying it.)

Gemini 2.0 Experimental Advanced says:

-------snip--------

This document is a petition filed in the Orphans Court of Baltimore County. Here's a breakdown of its content:

Petitioners: Neary Senters and his wife are petitioning the court.

Subject: The petition concerns a boy named James Gannon, who is identified as a pauper and has been bound (apprenticed) to them.

Issue: The petitioners claim that James Gannon is claimed by a certain Henry Haine of Baltimore City. They believe Haine is attempting to exploit the boy's labor. He obtained the boy through a seemingly valid indenture (apprenticeship contract) authorized by two Justices of the Peace in Baltimore City. However, the indenture references records in the Register of Wills' office in Baltimore City, suggesting the boy was already an indentured servant. The petitioners argue that a 1793 Maryland law (Chapter 45) regarding the binding of apprentices invalidates Haine's claim. They assert that the law requires parental consent for binding out a child, and there's no evidence of such consent in this case.

Request: The petitioners request the court to summon Henry Haine and the Justices of the Peace involved in creating the indenture. They also seek to have James Gannon and his parents (if living) appear before the court. Their ultimate goal is to have the indenture to Haine declared invalid, thus retaining custody and the labor of the boy.

Legal Argument: The core of their argument hinges on the 1793 Maryland law, which they believe was violated when James Gannon was bound to Henry Haine without parental consent. They are essentially challenging the legality of the indenture held by Haine.

In essence, this document reveals a dispute over the legal custody and labor of a young boy in 18th-century Baltimore. The petitioners are using a specific law about apprenticeship to argue against another individual's claim to the boy, highlighting the importance of legal procedures and parental rights in such matters during that time.

-------end snip--------

The prompt was simply "Please provide a detailed summary of this document" (I wasn't sure what else to ask.)

Is that closer to your wife's interpretation? I mean, I can't make head or tail of it...

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#426

Earlier quoted context omitted.

This happens literally every time. Someone always says "ChatGPT can't do this!", but then when someone actually runs the example, chatGPT gets it right. Now what the OP is going to do next is proceed to move goalposts and say like "but umm I just asked chatgpt this, so clearly they modified the code in realtime to get the answer right"

Similarly, in every thread there’s an AI skeptic who says LLMs are “useless” for coding, and never provides an example query for what they were trying.

Because the argument isn't based on individual query results. See for example my comment on a previous post https://news.ycombinator.com/item?id=42563715 .

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#427
I don't get why this matters at all? I looked at the o1-preview paper, Putnam is not mentioned. Meaning that: a) OpenAI never claimed this model achieves X% on this dataset. b) likely, OpenAI did not take measures to exclude this dataset from training. Meaning the only conclusion we can draw from this result is: when prompted with questions that were verbarim in the dataset, performance increases dramatically. We already know this, and it doesn't say anything about the performance of the model on unseen problems.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#428

Earlier quoted context omitted.

but e=mc^2 is just an approximation e: nice, downvoted for knowing special relativity

This thread has come up before(1), but I'll continue to argue that relativistic mass is a perfectly valid concept as much as any other, and if you disagree, you'll need arguments more substantial than it just being unpopular these days. Especially if you're trying to argue people out of using a concept that they personally find useful to aid their own understanding, just because it doesn't fit your own mathematical o…

sure but i think it

1. is not very intuitive/useful to have mass that varies on the direction (which is what this implies)

2. is somewhat tautological to define a new mass m_rel = E/c^2 and say that it satisfies the equation when this is not what most people understand mass to be. most people understand photons to be massless particles.

at minimum, relativistic mass should always be specified as m_rel to distinguish from what is typically referred to as mass.

but i don’t think relativistic mass is a wrong concept any more than any other mathematical convenience like virtual particles. the main question is how useful is it and should it be described using the word “mass” or is this confusing. there is value in having shared language, even if you can construct an alternate system of symbols and rules that can yield the same answer to every question. to the extent to which intent of the author matters at all (probably doesn’t), Einstein agreed that relativistic mass was not a useful concept.

i'll concede that the arguments in the thread you linked are not good

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#429
post #338

Earlier quoted context omitted.

But what does this have to do with reasoning? Yes, LLMs are not knowledge bases, and seeing people treat them as such absolutely terrifies me. However, I don’t see how the fact that LLMs often hallucinate “facts” is relevant to a discussion about their reasoning capabilities.

"Hallucinating a fact" that isn't in the training set and is also illogical, is exactly what a failure to reason correctly looks like.

Reasoning involves making accurate inferences based on the information provided in the current context, rather than recalling arbitrary facts from the training data.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#430
post #163

Earlier quoted context omitted.

https://s.h4x.club/bLuNed45 - it's more crazy to me that my wife CAN in fact read this stuff easily, vs the fact that an LLM can't. (for anyone who doesn't feel like downloading the zip, here is a single image from the zip: https://s.h4x.club/nOu485qx )

One stumbling block is that OpenAI's best model (o1-pro) can't yet deal with image uploads (edit: I'm wrong, it works now... trying it.) Gemini 2.0 Experimental Advanced says: -------snip-------- This document is a petition filed in the Orphans Court of Baltimore County. Here's a breakdown of its content: Petitioners: Neary Senters and his wife are petitioning the court. Subject: The petition concerns a boy named Jam…

Ideas are kinda there, details (like names) are wrong wrong.

Feel like I should make this task my personal ASI benchmark lol :D

Post reply on HN