Live data from Hacker News

Notes on OpenAI's new o1 chain-of-thought models

simonwillison.net

331–340 of 659 posts

Re: Notes on OpenAI's new o1 chain-of-thought models

#331
post #167

Earlier quoted context omitted.

What I'm not able to comprehend is why people are not seeing the answer as brilliant! Any ordinary mortal (like me) would have jumped to the conclusion that answer is "Father" and would have walked away patting on my back, without realising that I was biased by statistics. Whereas o1, at the very outset smelled out that it is a riddle - why would anyone out of blue ask such question. So, it started its chain of thoug…

The 'riddle': A woman and her son are in a car accident. The woman is sadly killed. The boy is rushed to hospital. When the doctor sees the boy he says "I can't operate on this child, he is my son". How is this possible? GPT Answer: The doctor is the boy's mother Real Answer: Boy = Son, Woman = Mother (and her son), Doctor = Father (he says...he is my son) This is not in fact a riddle (though presented as one) and th…

It literally is a riddle, just as the original one was, because it tries to use your expectations of the world against you. The entire point of the original, which a lot of people fell for, was to expose expectations of gender roles leading to a supposed contradiction that didn't exist.

You are now asking a modified question to a model that has seen the unmodified one millions of times. The model has an expectation of the answer, and the modified riddle uses that expectation to trick the model into seeing the question as something it isn't.

That's it. You can transform the problem into a slightly different variant and the model will trivially solve it.

Re: Notes on OpenAI's new o1 chain-of-thought models

#332
post #167

Earlier quoted context omitted.

What I'm not able to comprehend is why people are not seeing the answer as brilliant! Any ordinary mortal (like me) would have jumped to the conclusion that answer is "Father" and would have walked away patting on my back, without realising that I was biased by statistics. Whereas o1, at the very outset smelled out that it is a riddle - why would anyone out of blue ask such question. So, it started its chain of thoug…

The 'riddle': A woman and her son are in a car accident. The woman is sadly killed. The boy is rushed to hospital. When the doctor sees the boy he says "I can't operate on this child, he is my son". How is this possible? GPT Answer: The doctor is the boy's mother Real Answer: Boy = Son, Woman = Mother (and her son), Doctor = Father (he says...he is my son) This is not in fact a riddle (though presented as one) and th…

The original riddle is of course:

"A father and his son are in a car accident [...] When the boy is in hospital, the surgeon says: This is my child, I cannot operate on him".

In the original riddle the answer is that the surgeon is female and the boy's mother. The riddle was supposed to point out gender stereotypes.

So, as usual, ChatGPT fails to answer the modified riddle and gives the plagiarized stock answer and explanation to the original one. No intelligence here.

Re: Notes on OpenAI's new o1 chain-of-thought models

#333
post #167

Earlier quoted context omitted.

What I'm not able to comprehend is why people are not seeing the answer as brilliant! Any ordinary mortal (like me) would have jumped to the conclusion that answer is "Father" and would have walked away patting on my back, without realising that I was biased by statistics. Whereas o1, at the very outset smelled out that it is a riddle - why would anyone out of blue ask such question. So, it started its chain of thoug…

Come on. Of course chatgpt has read that riddle and the answer 1000 times already.

It hasn't read that riddle because it is a modified version. The model would in fact solve this trivially if it _didn't_ see the original in its training. That's the entire trick.

Re: Notes on OpenAI's new o1 chain-of-thought models

#334

Earlier quoted context omitted.

building SOTA systems is the easy part?! Easy compared to what?

Probably, to get them to work without hallucinating, or without failing a good percentage of the time.

I wonder what would our world look like if these two expectations that you seem to be taking for granted were applied to our politicians.

Re: Notes on OpenAI's new o1 chain-of-thought models

#335
post #51
post #18

The o1-preview model still hallucinates non-existing libraries and functions for me, and is quickly wrong about facts that aren't well-represented on the web. It's the usual string of "You're absolutely correct, and I apologize for the oversight in my previous response. [Let me make another guess.]" While the reasoning may have been improved, this doesn't solve the problem of the model having no way to assess if what…

I honestly can’t believe this is the hyped up “strawberry” everyone was claiming is pretty much AGI. Senior employees leaving due to its powers being so extreme I’m in the “probabilistic token generators aren’t intelligence” camp so I don’t actually believe in AGI, but I’ll be honest the never ending rumors / chatter almost got to me Remember, this is the model some media outlet reported recently that is so powerful…

"Senior employees leaving due to its powers being so extreme"

This never happened. No one said it happened.

"the model some media outlet reported recently that is so powerful OAI is considering charging $2k/month for"

The Information reported someone at a meeting suggested this for future models, not specifically Strawberry, and that it would probably not actually be that high.

Re: Notes on OpenAI's new o1 chain-of-thought models

#336

Earlier quoted context omitted.

The failure is in how you're using it. I don't mean this as a personal attack, but more to shed light on what's happening. A lot of people use LLMs as a search engine. It makes sense - it's basically a lossy compressed database of everything its ever read, and it generates output that is statistically likely - varying degrees of likeliness depending on the temperature, as well as how many times the particular weights…

> The failure is in how you're using it. People, for the most part, know what they know and don't know. I am not uncertain that the distance between the earth and the sun varies, but I'm certain that I don't know the distance from the earth to the sun, at least not with better precision than about a light week. This is going to have to be fixed somehow to progress past where we are now with LLMs. Maybe expecting an L…

> the distance from the earth to the sun, at least not with better precision than about a light week

The sun is eight light minutes away.

Re: Notes on OpenAI's new o1 chain-of-thought models

#337
post #283

I posted this on the other thread, but the two tests I had, it passed when ChatGPT-4 failed. https://chatgpt.com/share/66e35c37-60c4-8009-8cf9-8fe61f57d3... https://chatgpt.com/share/66e35f0e-6c98-8009-a128-e9ac677480...

The farmer riddle isn't quite right as you presented it. One of the parts that makes it interesting is that the boat can't carry everything at one time[1]. It can't happen in one trip; something must be left behind. It solved the correct version fine: https://chatgpt.com/share/66e3f9bb-632c-8005-9c95-142424e396... 1: https://en.wikipedia.org/wiki/Wolf,_goat_and_cabbage_problem

You misunderstand the situation.

If I give ChatGPT-4 the original farmer riddle, it "solves" it just fine, but it's assumed that it isn't actually solving it. That is, it's not thinking or doing any logical reasoning, or anything resembling that to come to a solution to the problem, but that it's simply regurgitating the problem's solution since it appears in the training data.

Giving ChatGPT-4 the modified farmers riddle, and having it spit out the incorrect, multi-step solution, is then proof that the LLM isn't doing anything that can be considered reasoning, but that it's merely repeating what's assumed to be in its training data.

ChatGPT-o1-preview correctly managing to actually parse my modified riddle, and then not simply parroting out the answer from the training corpus but give the right solution, as if it read it carefully, then says something about the improved logical and deductive reasoning capabilities of the newer model.

Re: Notes on OpenAI's new o1 chain-of-thought models

#338

I thought with this chain-of-thought approach the model might be better suited to solve a logic puzzle, e.g. ZebraPuzzles [0]. It produced a ton of "reasoning" tokens but hallucinated more than half of the solution with names/fields that weren't available. Not a systematic evaluation, but it seems like a degradation from 4o-mini. Perhaps it does better with code reasoning problems though -- these logic puzzles are es…

o1-mini does better than any other model on zebra puzzles. Maybe you got unlucky on one question?

https://www.reddit.com/r/LocalLLaMA/comments/1ffjb4q/prelimi...

Re: Notes on OpenAI's new o1 chain-of-thought models

#339

>> which is larger 3.9, 3.10 or 3.11 To determine which number is larger among 3.9, 3.10, and 3.11, let's compare Standard Decimal Interpretation: 3.9 is simply 3.9. 3.10 is equivalent to 3.1 because trailing zeros after the decimal point do not add value in decimal numbers. 3.11 remains 3.11. Comparison: 3.9 vs. 3.1 vs. 3.11 3.11 is the largest number. Version Number Interpretation: In version numbering, each segmen…

This is truly the new model's answer ? It's pretty similar to 3.5's "reasoning" actually:

In this context, "3.10" and "3.11" should be interpreted as decimal numbers, not as numbers with more digits.

When comparing:

3.9 3.10 (which is equal to 3.1) 3.11 (which is equal to 3.11) We have:

3.9 is greater than 3.1 (3.10), because 9 is larger than 1. 3.11 is greater than 3.9, because 11 is larger than 9. Thus, 3.11 is the largest of the three numbers.

Re: Notes on OpenAI's new o1 chain-of-thought models

#340

I am mostly only an LLM user with technical background. I don't have much in-depth knowledge. So I have questions about this take: >the output token allowance has been increased dramatically—to 32,768 for o1-preview and 65,536 for the supposedly smaller o1-mini! So the text says reasoning and output tokens are the same, as in you pay for both. But does the increase say that it can actually do more, or does it just me…

I included that note because output limits are a personal interest of mine.

Until recently most models capped out at around 4,000 tokens of output, even as they grew to handle 100,000 or even a million input tokens.

For most use-cases this is completely fine - but there are some edge-cases that I care about. One is translation - if you feed in a 100,000 token document in English and ask for it to be translated to German you want about 100,000 tokens of output, rather than a summary.

The second is structured data extraction: I like being able to feed in large quantities of unstructured text (or images) and get back structured JSON/CSV. This can be limited by low output token counts.

Post reply on HN