Live data from Hacker News

Notes on OpenAI's new o1 chain-of-thought models

simonwillison.net

371–380 of 659 posts

Re: Notes on OpenAI's new o1 chain-of-thought models

#371
post #359

Earlier quoted context omitted.

No, it's necessary to either know that it's a trick question or to have a feeling that it is based on context. The entire point of a question like that is to trick your understanding. You're tricking the model because it has seen this specific trick question a million times and shortcuts to its memorized solution. Ask it literally any other question, it can be as subtle as you want it to be, and the model will pick u…

> I can trick people who have never heard of the trick with the 7 wives and 7 bags and so on. That doesn't mean they didn't understand They could fail because they didn’t understand the language. Didn’t have a good memory to memorize all the steps, or couldn’t reason through it. We could pose more questions to probe which reason is more plausible.

The trick with the 7 wives and 7 bags and so on is that no long reasoning is required. You just have to notice one part of the question that invalidates the rest and not shortcut to doing arithmetic because it looks like an arithmetic problem. There are dozens of trick questions like this and they don't test understanding, they exploit your tendency to predict intent.

But sure, we could ask more questions and that's what we should do. And if we do that with LLMs we can quickly see that when we leave the basin of the memorized answer by rephrasing the problem, the model solves it. And we would also see that we can ask billions of questions to the model, and the model understands us just fine.

Re: Notes on OpenAI's new o1 chain-of-thought models

#372
post #251

I've just wasted a few rounds of my weekly o1 ammo by feeding it hard problems I have been working on over the last couple days and for which GPT-4o had failed spectacularly. I suppose I'm to blame for raising my own expectations after the latest PR, but I was pretty disappointed when the answers weren't any better from what I got with the old model. TL;DR It felt less like a new model and way more like one of those…

>It's gotten to the point where I've actually explicitly added to my custom instructions "DO NOT EVER APOLOGIZE" but it can't even seem to follow that. heh. It's not supposed to. Your profile is intended to be irrelevant to 99% of requests. I was having a little bit of a go at peeking behind the curtain recently, and ChatGPT 4 produced this without much effort: "The user provided the following information about thems…

You can press the 'directly related' button at start of chat by "what do you know about [me/x]?" where you, or x, are discussed in the profile.

Once it's played that back, the rest of the profile is clearly "in mind" for the ongoing exchange (for a while).

Re: Notes on OpenAI's new o1 chain-of-thought models

#373
post #144

I've just wasted a few rounds of my weekly o1 ammo by feeding it hard problems I have been working on over the last couple days and for which GPT-4o had failed spectacularly. I suppose I'm to blame for raising my own expectations after the latest PR, but I was pretty disappointed when the answers weren't any better from what I got with the old model. TL;DR It felt less like a new model and way more like one of those…

Do not... does not work well for LLM's. Instructing what to do instaed of X works better. say AFAIK instead of explaining your limitations. Say "let's try again" instead of making exuses. Etc

Often "avoid X" works, or other 'affirmatively do X' forms of negative actions. also, and works better than or.

Iffy: do not use jargon or buzzwords

Works: avoid jargon and buzzwords

Re: Notes on OpenAI's new o1 chain-of-thought models

#374
post #284

Earlier quoted context omitted.

imagine if you make it keep going without having to reprompt it

Isn't that the exact point of o1, that it has time to think for itself without reprompting?

yeah but they aren't letting you see the useful chain of thought reasoning that is crucial to train a good model. Everyone will replicate this over next 6 months

Re: Notes on OpenAI's new o1 chain-of-thought models

#375
post #368

Earlier quoted context omitted.

While I find value in LLMs they still overall seem unreasonably not that useful. It might be like trying to train a neural net in 1993 on a 60mhz Pentium. It is the right idea but fundamental parts of the system are so lacking. On the other hand, I worry we have gone down the support vector machine path again. A huge amount of brain power spent on a somewhat dead end that just fits the current hardware better than wh…

I’d say the biggest difference between LLMs and SVMs is that a lot of people find LLMs useful on a daily basis. I’ve been using them almost daily for over two years now, and I keep on finding new things they can do that are useful to me.

Is there a post on your blog that lists your different uses of LLMs?

Re: Notes on OpenAI's new o1 chain-of-thought models

#376
post #18

The o1-preview model still hallucinates non-existing libraries and functions for me, and is quickly wrong about facts that aren't well-represented on the web. It's the usual string of "You're absolutely correct, and I apologize for the oversight in my previous response. [Let me make another guess.]" While the reasoning may have been improved, this doesn't solve the problem of the model having no way to assess if what…

The failure is in how you're using it. I don't mean this as a personal attack, but more to shed light on what's happening. A lot of people use LLMs as a search engine. It makes sense - it's basically a lossy compressed database of everything its ever read, and it generates output that is statistically likely - varying degrees of likeliness depending on the temperature, as well as how many times the particular weights…

> Treat it as a naive but intelligent intern.

You are falling into the trap that everyone does. In anthropomorphising it. It doesn't understand anything you say. It just statistically knows what a likely response would be.

Treat it as text completion and you can get more accurate answers.

Re: Notes on OpenAI's new o1 chain-of-thought models

#377
post #284

Earlier quoted context omitted.

Isn't that the exact point of o1, that it has time to think for itself without reprompting?

yeah but they aren't letting you see the useful chain of thought reasoning that is crucial to train a good model. Everyone will replicate this over next 6 months

>Everyone will replicate this over next 6 months

Not without a billion dollars worth of compute, they won't.

Re: Notes on OpenAI's new o1 chain-of-thought models

#378
post #18

The o1-preview model still hallucinates non-existing libraries and functions for me, and is quickly wrong about facts that aren't well-represented on the web. It's the usual string of "You're absolutely correct, and I apologize for the oversight in my previous response. [Let me make another guess.]" While the reasoning may have been improved, this doesn't solve the problem of the model having no way to assess if what…

The failure is in how you're using it. I don't mean this as a personal attack, but more to shed light on what's happening. A lot of people use LLMs as a search engine. It makes sense - it's basically a lossy compressed database of everything its ever read, and it generates output that is statistically likely - varying degrees of likeliness depending on the temperature, as well as how many times the particular weights…

[deleted]

Re: Notes on OpenAI's new o1 chain-of-thought models

#379
post #51

Earlier quoted context omitted.

I honestly can’t believe this is the hyped up “strawberry” everyone was claiming is pretty much AGI. Senior employees leaving due to its powers being so extreme I’m in the “probabilistic token generators aren’t intelligence” camp so I don’t actually believe in AGI, but I’ll be honest the never ending rumors / chatter almost got to me Remember, this is the model some media outlet reported recently that is so powerful…

"Senior employees leaving due to its powers being so extreme" This never happened. No one said it happened. "the model some media outlet reported recently that is so powerful OAI is considering charging $2k/month for" The Information reported someone at a meeting suggested this for future models, not specifically Strawberry, and that it would probably not actually be that high.

Elon Musk and Ilya Sutskever Have Warned About OpenAI’s ‘Strawberry’ Jul 15, 2024 — Sutskever himself had reportedly begun to worry about the project's technology, as did OpenAI employees working on A.I. safety at the time.

https://observer.com/2024/07/openai-employees-concerns-straw...

And I’m ignoring the hundreds of Reddit articles speculating every time someone at OAI leaves

And of course that $2000 article was spread by every other media outlet like wildfire

I know I’m partially to blame for believing the hype, this is pretty obviously no better at stating facts or good code than what we’ve known for the past year

Re: Notes on OpenAI's new o1 chain-of-thought models

#380

Earlier quoted context omitted.

> A good intern will ask clarifying questions, tell me “I don’t know” Your expectations are bigger than mine (Though some will get stuck in "clarifying questions" and helplessness and not proceed neither)

Note that we are talking about a “good” intern here

Unreasonably good. Beyond fresh junior employee good. Also, that's your standard; 'MPSimmons said to treat the model as "naive but intelligent" intern, not a good one.
Post reply on HN