Live data from Hacker News

30% drop in O1-preview accuracy when Putnam problems are slightly variated

openreview.net

431–440 of 558 posts

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#431
post #420

Earlier quoted context omitted.

I must be to tired as I can't find any flaw in that reasoning.

The joke/riddle text is "he says" but Claude says "their son" and suggests the doctor might be a woman. More substantively: "This is a classic riddle that highlights gender bias - many people assume doctors must be men, but don't initially consider that the doctor could be the father." is totally nonsensical. The text is a gender (and meaning) inversion of the classic riddle to confuse LLMs. Even though Claude correc…

[deleted]

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#432

Earlier quoted context omitted.

I don't know why I and you are getting downvoted. Sometimes, HN crowd is just unhinged against AI.

These models are trained in two steps: training base model and then uptraining it. First step includes as much data as possible, everything company can find. For Llama models it's 15T tokens, which is ~40 TB of data. No-one really puts an effort on splitting this data into train/test/eval (and it's not very achievable either). It's just as much data as possible. So it's like 99.9999999% wrong to assume something publ…

right, but where did someone assume it wasn’t in the train set? they just said it wasn’t in the test set

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#433

Earlier quoted context omitted.

Okay, then what about elite level codeforces performance? Those problems weren’t even constructed until after the model was made. The real problem with all of these theories is most of these benchmarks were constructed after their training dataset cutoff points. A sudden performance improvement on a new model release is not suspicious. Any model release that is much better than a previous one is going to be a “sudden…

o1 has a ~1650 rating, at that level many or most problems you will be solving are going to be a transplant of a relatively known problem. Since o1 on codeforces just tried hundreds or thousands of solutions, it's not surprising it can solve problems where it is really about finding a relatively simple correspondence to a known problem and regurgitating an algorithm. In fact when you run o1 on ""non-standard"" codefo…

o3 is what i’m referring to and it is 2700

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#434
post #430

Earlier quoted context omitted.

One stumbling block is that OpenAI's best model (o1-pro) can't yet deal with image uploads (edit: I'm wrong, it works now... trying it.) Gemini 2.0 Experimental Advanced says: -------snip-------- This document is a petition filed in the Orphans Court of Baltimore County. Here's a breakdown of its content: Petitioners: Neary Senters and his wife are petitioning the court. Subject: The petition concerns a boy named Jam…

Ideas are kinda there, details (like names) are wrong wrong. Feel like I should make this task my personal ASI benchmark lol :D

o1-pro just answered (it often takes a few minutes). Not sure if this is any better (is it?) but at least it is starting to admit when it isn't sure about something, rather than just spewing BS confidently. That seems like a genuine win.

---------snip--------

    From the handwriting and context, this appears to be a mid-19th-century petition submitted to the Orphans’ Court of Baltimore County. Below is a point‐by‐point summary of its content and purpose:

    Address and Purpose
        The document is formally addressed “To the Honorable Judges of the Orphans’ Court for Baltimore County.”
        It identifies itself as a petition—that is, a request for the Court’s action or permission under the relevant orphan or apprenticeship laws.

    Parties Involved
        The petitioners are (1) Henry Jenkins and (2) another individual whose name is difficult to read, but who appears to be the child’s mother (or at least a close guardian figure).
        They describe themselves as residents of Baltimore City and ask the Court’s intervention concerning a minor boy—also named “Henry” (it looks like “Henry Thomas,” though the middle/family name is partly obscured).

    Relationship to the Minor
        The petition states that the mother had this boy by a former husband, meaning the child is not the biological son of Henry Jenkins.
        The boy is described as being about fifteen (or so) years of age.
        The petitioners refer to him as an “orphan,” suggesting his father is deceased and that they (the petitioners) are seeking legal authority over him, or wish the court to confirm or formalize that authority.

    Legal Concern: Binding as an Apprentice
        A key theme is the binding of a minor as an apprentice under Maryland law. In early Baltimore history (and in many states at that time), orphans or children whose parents were unable to support them could be legally “bound” (apprenticed) so that the child would learn a trade and have financial support.
        The petition cites the Act of 1788, Chapter 45, which was one of Maryland’s statutes governing apprenticeships—how they are created, what formalities must be met, and how the Orphans’ Court is involved.

    Specifics of the Request
        The petitioners believe there is either a gap or question about whether they have the legal right to apprentice the boy. They are asking the Orphans’ Court to recognize or approve the arrangement.
        They indicate that the mother (as the surviving parent) and Henry Jenkins (presumably the stepfather or the mother’s new husband) want to place the boy under an apprenticeship, possibly so he can learn a trade and be supported until he reaches the age of majority.

    References to Official Records
        The document mentions the Register of Wills and/or Deeds in Baltimore City, implying that they have looked into or recorded something about the child’s status.
        They point out that the legal forms required for binding a child as an apprentice often must be recorded with an official clerk or the Orphans’ Court, and they want to ensure everything is in proper order.

    Legal Justification
        The petitioners note that under the relevant laws, it “is lawful to bind out any child” whose parents cannot fully support them, or when the father is deceased.
        They appear to lay out the reasons why the boy qualifies for such an arrangement (i.e., he is fatherless and his mother/stepfather lack the resources to maintain him or want to formalize his situation).

    Request for Court Action
        Ultimately, they are asking the Court to review the circumstances and grant them permission (or confirm their right) to bind the boy as an apprentice.
        They emphasize that the law requires the Orphans’ Court to oversee such arrangements to protect the minor’s welfare and ensure the terms are fair.

    Overall Significance:
       This document is a formal petition under Maryland’s 18th- and 19th-century apprenticeship laws. It centers on a minor boy (about fifteen years old), the remarriage of his mother, and the need to have the Orphans’ Court recognize or approve an apprenticeship arrangement. The purpose is to comply with the legal requirements of binding a minor so that he can be taught a trade and maintained financially until adulthood.
---------end snip--------

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#435
post #197

I remember when this stuff was all coming out and people were finally excited about ChatGPT getting the problem with "which is heavier, a 10 pound bag of feathers or a 10 pound bag of bricks?" problem correct. But of course it got it correct. It was in the training set. Vary the problem slightly by just changing the nouns, or changing the numbers so that one in fact was heavier than the other, and performance went al…

o1 crushes your variant. Here's the response.

Step-by-Step Reasoning 1. Clarify what is being compared The question asks us to compare two bags: one containing steel ingots weighing 9.99 pounds and another containing fluffy cotton weighing 10.01 pounds. 2. Check the actual weight Despite our intuition that steel is “heavier” than cotton by density, the problem specifically states the total weight in pounds of each bag. Since 9.99 Answer

A 10.01-pound bag of fluffy cotton is heavier than a 9.99-pound bag of steel ingots, because 10.01 pounds is strictly greater than 9.99 pounds.

Reference • Basic arithmetic: 10.01 is greater than 9.99. • For a playful twist on a similar concept, see any version of the riddle “What weighs more—a pound of feathers or a pound of lead?” In that classic riddle, both weigh the same; here, the numbers differ.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#436

Earlier quoted context omitted.

These models are trained in two steps: training base model and then uptraining it. First step includes as much data as possible, everything company can find. For Llama models it's 15T tokens, which is ~40 TB of data. No-one really puts an effort on splitting this data into train/test/eval (and it's not very achievable either). It's just as much data as possible. So it's like 99.9999999% wrong to assume something publ…

right, but where did someone assume it wasn’t in the train set? they just said it wasn’t in the test set

What test set is being talked about here? Why does it matter what’s on this set?

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#437
post #427

I don't get why this matters at all? I looked at the o1-preview paper, Putnam is not mentioned. Meaning that: a) OpenAI never claimed this model achieves X% on this dataset. b) likely, OpenAI did not take measures to exclude this dataset from training. Meaning the only conclusion we can draw from this result is: when prompted with questions that were verbarim in the dataset, performance increases dramatically. We alr…

yep, welcome to hn

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#438
post #197

I remember when this stuff was all coming out and people were finally excited about ChatGPT getting the problem with "which is heavier, a 10 pound bag of feathers or a 10 pound bag of bricks?" problem correct. But of course it got it correct. It was in the training set. Vary the problem slightly by just changing the nouns, or changing the numbers so that one in fact was heavier than the other, and performance went al…

ChatGPT Plus user here. The following are all fresh sessions and first answers, no fishing. GPT 4: The 10.01-pound bag of fluffy cotton is heavier than the 9.99-pound bag of steel ingots. The type of material doesn’t affect the weight comparison; it’s purely a matter of which bag weighs more on the scale. GPT 4o: The 10.01-pound bag of fluffy cotton is heavier. Weight is independent of the material, so the bag of cot…

The question isn’t whether or not it can get this question correct or not. It is, why is it incapable of getting the answer consistently right?

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#439

Earlier quoted context omitted.

right, but where did someone assume it wasn’t in the train set? they just said it wasn’t in the test set

What test set is being talked about here? Why does it matter what’s on this set?

the point is that putnam was never a test/benchmark being used by OpenAI or anyone else, so there is no smoking gun if you find putnam on the train set nor is it cheating or nefarious because nobody ever claimed otherwise.

this whole notion of putnam as test being trained on is a fully invented grievance

read the entire thread in this context

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#440
post #64

Earlier quoted context omitted.

If it were a complete failure on variations I would be inclined to agree. Instead it was a 30% drop in performance. I would characterise that as limited understanding.

My guess is that what’s understood isn’t various parts of solving the problem but various aspects of the expected response. I see this more akin to a human faking their way through a conversation.

I see this more akin to a human faking their way through a conversation.

That works in English class. Try it in a math class and you'll get a much lower grade than ChatGPT will.

Post reply on HN