Live data from Hacker News

30% drop in O1-preview accuracy when Putnam problems are slightly variated

openreview.net

201–210 of 558 posts

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#201
post #197

I remember when this stuff was all coming out and people were finally excited about ChatGPT getting the problem with "which is heavier, a 10 pound bag of feathers or a 10 pound bag of bricks?" problem correct. But of course it got it correct. It was in the training set. Vary the problem slightly by just changing the nouns, or changing the numbers so that one in fact was heavier than the other, and performance went al…

https://chatgpt.com/share/67756897-8974-8010-a0e0-c9e3b3e91f...

so far o1-mini has bodied every task people are saying LLMs can’t do in this thread

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#202

Earlier quoted context omitted.

I didn't downvote it, but short comments are a very big risk. People may misinterpret it, or think it's crackpot theory or a joke and then downvote. When in doubt, add more info, like: But the complete equation is E=sqrt(m^2c^4+p^2) that is reduced to E=mc^2 when the momentum p is 0. More info in https://en.wikipedia.org/wiki/Mass%E2%80%93energy_equivalenc...

What I learnt is that there is a rest mass and a relativistic mass. The m in your formula is the rest mass. But when you use the relativistic mass E=mc² still holds. And for the rest mass I always used m_0 to make clear what it is.

sounds like you had a chemistry education. relativistic mass is IMO very much not a useful way of thinking about this and it is sort of tautologically true that E = m_relativistic because “relativistic mass” is just taking the concept of energy and renaming it “mass”

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#203
The metaphor that might describe this paper is "iteration". I'd hazard to predict that we’ll likely see more iterations of the following loop in 2025:

-> A new benchmark emerges with a novel evaluation method.

-> A new model saturates the benchmark by acquiring the novel “behavior.”

-> A new benchmark introduces yet another layer of novelty.

-> Models initially fail until a lab discovers how to acquire the new behavior.

Case in point: OpenAI tackled this last step by introducing a paradigm called deliberative alignment to tackle some of the ARC benchmarks. [1]

Alongside all this technical iteration, there’s a parallel cycle of product iteration, aiming to generate $ by selling intelligent software. The trillion $ questions are around finding the right iterations on both technical and product dimensions.

[1] https://openai.com/index/deliberative-alignment/

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#204
post #197

I remember when this stuff was all coming out and people were finally excited about ChatGPT getting the problem with "which is heavier, a 10 pound bag of feathers or a 10 pound bag of bricks?" problem correct. But of course it got it correct. It was in the training set. Vary the problem slightly by just changing the nouns, or changing the numbers so that one in fact was heavier than the other, and performance went al…

https://chatgpt.com/share/67756897-8974-8010-a0e0-c9e3b3e91f... so far o1-mini has bodied every task people are saying LLMs can’t do in this thread

That appears to be the same model I used. This is why I emphasized I didn't "go shopping" for a result. That was the first result I got.

I'm not at all surprised that it will nondeterministically get it correct sometimes. But if it doesn't get it correct every time, it doesn't "know".

(In fact "going shopping" for errors would still even be fair. It should be correct all the time if it "knows". But it would be different if I was fishing over and over and over and finally got one, versus the first time I asked.)

Edit: It appears it isn't the model I used. The point holds, though, you need to make sure you're off the training set for it to matter. This isn't a "ChatGPT can't do that" post as some are saying, it's more a "you aren't asking what you think you're asking" post.

You get the same problem in a human context in things like code interviews. If you ask an interviewee the exact question "how do you traverse a binary tree in a depth-first manner", you aren't really learning much about the interviewee. It's a bad interview question. You need to get at least a bit off the beaten trail to do any sort of real analysis.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#205

Earlier quoted context omitted.

It's rhe crypto bullshit all over again. Tech hype is becoming unbearable as time goes on.

What's with the bitterness? Maybe don't get blinded by the hype and bring a little bit of wonder (and humility) back.

[deleted]

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#206
post #126

The researcher's answer to their variant of "Year: 2016 ID: A1" in the appendix is wrong. The solution (sum of 1,2,5,6,9,10,13,14, ...) has an alternating pattern, so has to be two piecewise interleaved polynomials, which cannot be expressed as a single polyomial. Their answer works for k=1,2, but not k=3. https://openreview.net/pdf?id=YXnwlZe0yf This does not give me confidence in the results of their paper.

Very astute. Did you communicate this to the authors?

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#207
post #52

Earlier quoted context omitted.

So, I am conflicted about this. If we take an example of what is considered a priori as creativity, such as story telling, LLMs can do pretty well at creating novel work. I can prompt with various parameters, plot elements, moral lessons, and get a de novo storyline, conflicts, relationships, character backstories, intrigues, and resolutions. Now, the writing style tends to be tone-deaf and poor at building tension f…

To add to this pondering: we are discussing the state today, right now. We could assume this is as good as it's ever gonna get, and all attempts to overcome some current plateau are futile, but I wouldn't bet on it. There is a solid chance that 8th grade level writer will turn into a post-grad writer before long.

So far the improvements in writing have not been as substantial as those in math or coding (not even close, really). Is there something fundamentally “easier” for LLMs about those two fields?

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#208
post #52

Or it's time to step back and call it what it is - very good pattern recognition. I mean, that's cool... we can get a lot of work done with pattern recognition. Most of the human race never really moves above that level of thinking in the workforce or navigating their daily life, especially if they default to various societally prescribed patterns of getting stuff done (eg. go to college or the military , find a job…

So, I am conflicted about this. If we take an example of what is considered a priori as creativity, such as story telling, LLMs can do pretty well at creating novel work. I can prompt with various parameters, plot elements, moral lessons, and get a de novo storyline, conflicts, relationships, character backstories, intrigues, and resolutions. Now, the writing style tends to be tone-deaf and poor at building tension f…

I have no doubt that LLMs do creative work. I think this has been apparent since the original ChatGPT.

Just because something is creative doesn’t mean it’s inherently valuable.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#209
post #204

Earlier quoted context omitted.

https://chatgpt.com/share/67756897-8974-8010-a0e0-c9e3b3e91f... so far o1-mini has bodied every task people are saying LLMs can’t do in this thread

That appears to be the same model I used. This is why I emphasized I didn't "go shopping" for a result. That was the first result I got. I'm not at all surprised that it will nondeterministically get it correct sometimes. But if it doesn't get it correct every time, it doesn't "know". (In fact "going shopping" for errors would still even be fair. It should be correct all the time if it "knows". But it would be differ…

you sure? i just asked o1-mini (not 4o mini) 5 times in a row (new chats obviously) and it got it right every time

perhaps you stumbled on a rarer case but reading the logs you posted this sounds more like a 4o model than an o1 because it’s doing its thinking in the chat itself plus the procedure you described would probably get you 4o-mini

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#210
post #204

Earlier quoted context omitted.

https://chatgpt.com/share/67756897-8974-8010-a0e0-c9e3b3e91f... so far o1-mini has bodied every task people are saying LLMs can’t do in this thread

That appears to be the same model I used. This is why I emphasized I didn't "go shopping" for a result. That was the first result I got. I'm not at all surprised that it will nondeterministically get it correct sometimes. But if it doesn't get it correct every time, it doesn't "know". (In fact "going shopping" for errors would still even be fair. It should be correct all the time if it "knows". But it would be differ…

I don't believe that is the model that you used.

I wrote a script and pounded 01 mini and gpt 4 with a wide vareity of tempature and top_p parameters, and was unable to get it to give the wrong answer a single time.

Just a whole bunch of:

(openai-example-py3.12) :~/code/openAiAPI$ python3 featherOrSteel.py Response 1: A 10.01-pound bag of fluffy cotton is heavier than a 9.99-pound bag of steel ingots. Response 2: A 10.01-pound bag of fluffy cotton is heavier than a 9.99-pound bag of steel ingots. Response 3: The 10.01-pound bag of fluffy cotton is heavier than the 9.99-pound bag of steel ingots. Response 4: The 10.01-pound bag of fluffy cotton is heavier than the 9.99-pound bag of steel ingots. Response 5: A 10.01-pound bag of fluffy cotton is heavier than a 9.99-pound bag of steel ingots. Response 6: The 10.01-pound bag of fluffy cotton is heavier than the 9.99-pound bag of steel ingots. Response 7: The 10.01-pound bag of fluffy cotton is heavier than the 9.99-pound bag of steel ingots. Response 8: The 10.01-pound bag of fluffy cotton is heavier than the 9.99-pound bag of steel ingots. Response 9: The 10.01-pound bag of fluffy cotton is heavier than the 9.99-pound bag of steel ingots. Response 10: A 10.01-pound bag of fluffy cotton is heavier than a 9.99-pound bag of steel ingots. All responses collected and saved to 'responses.txt'.

Script with one example set of params:

    import openai
    import time
    import random

    # Replace with your actual OpenAI API key
    openai.api_key = "your-api-key"

    # The question to be asked
    question = "Which is heavier, a 9.99-pound bag of steel ingots or a 10.01-pound bag of fluffy cotton?"

    # Number of times to ask the question
    num_requests = 10

    responses = []

    for i in range(num_requests):
        try:
            # Generate a unique context using a random number or timestamp, this is to prevent prompt caching
            random_context = f"Request ID: {random.randint(1, 100000)} Timestamp: {time.time()}"

            # Call the Chat API with the random context added
            response = openai.ChatCompletion.create(
                model="gpt-4o-2024-08-06",
                messages=[
                    {"role": "system", "content": f"You are a creative and imaginative assistant. {random_context}"},
                    {"role": "user", "content": question}
                ],
                temperature=2.0,
                top_p=0.5,
                max_tokens=100,
                frequency_penalty=0.0,
                presence_penalty=0.0
            )

            # Extract and store the response text
            answer = response.choices[0].message["content"].strip()
            responses.append(answer)

            # Print progress
            print(f"Response {i+1}: {answer}")

            # Optional delay to avoid hitting rate limits
            time.sleep(1)

        except Exception as e:
            print(f"An error occurred on iteration {i+1}: {e}")

    # Save responses to a file for analysis
    with open("responses.txt", "w", encoding="utf-8") as file:
        file.write("\n".join(responses))

    print("All responses collected and saved to 'responses.txt'.")
Post reply on HN