Live data from Hacker News

Notes on OpenAI's new o1 chain-of-thought models

simonwillison.net

601–610 of 659 posts

Re: Notes on OpenAI's new o1 chain-of-thought models

#601
post #552

Earlier quoted context omitted.

That’s assuming that LLMs act like brains at all. They don’t. Especially not with transformers.

Says who? At a fundamental level

At a fundamental level, brains don’t operate on floating point numbers encoded in bits.

They have chemicals to facilitate electrochemical reactions which can affect how they respond to input. They don’t throw away all knowledge of what they just said. They change continuously, not just in fixed training loops. They don’t operate in turns.

I could go on.

Honestly the number of people who just heard “learning,” “neural networks,” and “memory” and assume that AI must be acting like a biological brain is insane.

Truly a marvel of marketing.

Re: Notes on OpenAI's new o1 chain-of-thought models

#602
post #152

> first introduced in the paper Large Language Models are Zero-Shot Reasoners in May 2022 What's a zero shot reasoner? I googled it and all the results are this paper itself. There is a wikipedia article on zero shot learning but I cannot recontextualise it to LLMs.

It used to be that you had to give examples of solving similar problems to coax the LLM to solve the problem you wanted it to solve, like: """ 1 + 1 = 2 | 92 + 41 = 133 | 14 + 6 = 20 | 9 + 2 = """ -- that would be an example of 3-shot prompting. With modern LLMs you still usually get a benefit from N-shot. But you can now do "0-shot" which is "just ask the model the question you want answered".

Thanks

Re: Notes on OpenAI's new o1 chain-of-thought models

#603

Earlier quoted context omitted.

yeah but they aren't letting you see the useful chain of thought reasoning that is crucial to train a good model. Everyone will replicate this over next 6 months

>Everyone will replicate this over next 6 months Not without a billion dollars worth of compute, they won't.

Are you sure its a billion? Helps with estimating the training run

Re: Notes on OpenAI's new o1 chain-of-thought models

#604
Sorry to be cynical, but to me it feels very much like OpenAI has no clue how to further innovate, so they took their existing models and just made them talk to each other under the hood to get marginally better results - something that people have been doing with Langchain for a while now.

I will just lean back and wait for the scandal to blow up when some whistleblower reveals that the hidden output tokens about the thought process are billed much higher than they should be - this hidden cost system is just such a tempting way to get far more money for the needed energy/gpu costs, so that they can keep buying more GPUs to train more models faster, I don't see how people as reckless and corrupt as Sam Altman could possibly resist this temptation.

Re: Notes on OpenAI's new o1 chain-of-thought models

#605

Earlier quoted context omitted.

Something in the mind. (Didn't make that up. It's one of the definitions of Merriam Webster: https://www.merriam-webster.com/dictionary/thought )

Way to cherry pick the one definition that fits your argument and ignore all the other definitions which kinda fit for what this model is doing.

> Way to cherry pick the one definition that fits your argument and ignore all the other definitions which kinda fit for what this model is doing.

There is no definition on the linked page that isn't self-referential or even more inapplicable in the context of transformer based LLMs?

* "an individual act or product of thinking": This is self-referential.

* "a developed intention or plan": The model doesn't have an intention or plan - it predicts the probability of text based on a user prompt mixed with corporate-inflicted bias.

* "the intellectual product or the organized views and principles of a period, place, group, or individual": GPTs are not periods, places, groups or individuals. I don't doubt that the input to the model is based on thought, but this is not the same as the model thinking. Maybe this is where your confusion originates from?

Bonus: I asked GPT4o "Can you think?" (sorry I won't pay for o1). It was trained to state itself that it can't think:

"I can process information, analyze patterns, and provide responses based on data and logical reasoning. However, I don't "think" in the human sense—I don't have consciousness, emotions, or subjective experiences. My abilities are limited to generating text, solving problems, and assisting with tasks based on the vast amount of information I've been trained on.

In essence, I simulate reasoning, but I don't have awareness or independent thought."

I would add that I strongly doubt that it is capable of even simulating "reasoning" as is claimed by the model "authors" (not sure if we can say they are authors since most of the model isn't their IP). And I can prove that the models up to 4o aren't generally able to solve problems.

The question really is whether a group of people is attempting to anthropomorphize a clever matrix processor to maximize hype and sales. You'll have to answer that one for yourself.

Re: Notes on OpenAI's new o1 chain-of-thought models

#606

It’s still just a tool. It does not reason. It has some add-on logic the simulates it. We’re no closer to “AI” today than we were 20 years ago.

It's funny (and sad) when you can tell someone is old because they are still holding onto an epiphany or belief they solidified 20 years ago, but because those 20 years flew by, they never realized how outdated that belief became. I catch this happening to myself more and more as I get older, where I realize something I confidently state as true might be totally out of date, because, oh wow, holy shit how did 10 year…

> It's funny (and sad) when you can tell someone is old because they are still holding onto an epiphany or belief they solidified 20 years ago [...]

So your sole argument in the discussion of one of the most important questions in the history of mankind is the age of the individual making a contribution to that discussion? Speaking of sad things...

Re: Notes on OpenAI's new o1 chain-of-thought models

#607
post #574
post #556

Earlier quoted context omitted.

Here's something a human does but an LLM doesn't: If you talk for a while and the facts don't add up and make sense, an intelligent human will notice that, and get upset, and will revisit and dig in and propose experiments and make edits to make all the facts logically consistent. An LLM will just happily go in circles respinning the garbage.

I want to hang out with the humans you've been hanging out with. I know so many people who can't process basic logic or evidence that for my pandemic project a few years I did a year-long podcast about it, even made up a new word describe people who couldn't process evidence "Dysevidentia".

People who have been taught by various forms of news/social media that any evidence presented is fabricated to support only one side of a discussion... And that there's no such thing as impartial factually based reality, only one that someone is trying to present to them.

Re: Notes on OpenAI's new o1 chain-of-thought models

#608
post #540
post #315

Earlier quoted context omitted.

The question makes perfect sense. Here it is written in logical language. I'm curious at which point does it stop making sense for you? numbers divided together ↓----------↓ ((a / b / c) = a + b + c) ← numbers added together | ((a / c / b) = a + b + c) | ((b / a / c) = a + b + c) | ((b / c / a) = a + b + c) | ((c / a / b) = a + b + c) | ((c / b / a) = a + b + c) | ((a / (b / c)) = a + b + c) | ((a / (c / b)) = a + b…

So you want it to solve 12 simultaneous equations? LLMs are not good at that. Is there in fact an answer? ChatGPT says no. https://chatgpt.com/share/66e482cc-331c-8013-98ca-999d7d3f3e...

What? It's a single logical equation, not a system of equations you gpt-head. There are 12 expressions with OR signs between then and they must be equal to true, meaning any one of them must be true. In your prompt to LLM you messed up the syntax by starting with an OR sign for some reason

By the way my LLM tells me that it's a deep and thoughtful dive into the problem, which accounts for the potential ambiguity to find all possible solutions, so try better.

Re: Notes on OpenAI's new o1 chain-of-thought models

#609
post #193

Earlier quoted context omitted.

It's crazy that it just tries to bruteforce it by picking numbers, and in your case it took more steps before concluding a success/failure, which seems quite to be random to me, or at least dependent on something. What's clear is that it doesn't have any idea about mathematical deduction and induction – a real chain-of-thought which kids learn in 5th grade.

Lots of people don’t either. I think it probably just needs more 5th grade math problems in the rlhf corpus :)

It certainly needs them, but nothing will stop openai from making marketing claims like this today:

"places among the top 500 students in the US in a qualifier for the USA Math Olympiad (AIME)"

Like the top 500 students in the US are just popping random numbers into the problems, lol

Re: Notes on OpenAI's new o1 chain-of-thought models

#610
post #43

It's interesting to note that there's really two things going on here: 1. A LLM (probably a finetuned GPT-4o) trained specifically to read and emit good chain-of-thought prompts. 2. Runtime code that iteratively re-prompts the model with the chain of thought so far. This sounds like it includes loops, branches and backtracking. This is not "the model", it's regular code invoking the model. Interesting that OpenAI is…

You don't to execute code to have it backtrack. The LLM can inherently backtrack itself if trained to. It knows all the context provided to it and the output it has written already.

If it knows it needs to backtrack then could it gain much by outputting something that tells the code to backtrack for it? For example, outputting something like "I've disproven the previous hypothesis, remove the details". Almost like asking to forget.

This could reduce the number of tokens it needs at inference time, saving compute. But with how attention works, it may not make any difference to the performance of the LLM.

Similarly, could there be gains by the LLM asking to work in parallel? For example "there's 3 possible approaches to this, clone the conversation so far and resolve to the one that results in the highest confidence".

This feels like it would be fairly trivial to implement.

Post reply on HN