Live data from Hacker News

Notes on OpenAI's new o1 chain-of-thought models

simonwillison.net

401–410 of 659 posts

Re: Notes on OpenAI's new o1 chain-of-thought models

#401
post #281

Near the end, the quote from OpenAI researcher Jason Wei seems damning to me: > Results on AIME and GPQA are really strong, but that doesn’t necessarily translate to something that a user can feel. Even as someone working in science, it’s not easy to find the slice of prompts where GPT-4o fails, o1 does well, and I can grade the answer. But when you do find such prompts, o1 feels totally magical. We all need to find…

He's speaking about his objective to make ever stronger LLMs: so for this his secondary objective is to measure their real performance. The human preference is not that good of a proxy measurement: for instance, it can be gamed by making the model more assertive, causing the human error-spotting ability to decrease a lot [0]. So what he's really saying is that non-rigorous human vibe checks (like those LMSys Chatbot…

ie when you cant beat them, make new metrics

and you can absolutely evaluate how smart someone is in a 2min casual conversation. You wont be able to tell how well they are in some niche topic, but %insert something about different flavors of intelligence and how they do not equate do subject matter expertise%

Re: Notes on OpenAI's new o1 chain-of-thought models

#403
post #47

Earlier quoted context omitted.

How many hours without it?

Not the OP, but in my experience LLMs fail in ways that indicate they will never solve the problem. Stuck in loops, correct their mistakes with worse mistakes, hallucinating things that don’t exist and being unable to correct. Working on my own, I have the confidence that I know I can make incremental forward progress on a problem. That’s much preferable.

But when working with an LLM you can still contribute.

Re: Notes on OpenAI's new o1 chain-of-thought models

#404
post #307
post #131

Earlier quoted context omitted.

> probabilistic token generators aren’t intelligence Maybe this has been extensively discussed before, but since I've lived under a rock: which parts of intelligence do you think are not representable as conditional probability distributions?

> which parts of intelligence do you think are not representable as conditional probability distributions Maybe I'm wrong here but a lot of our brilliance comes from acting against the statistical consensus. What I mean is, Nicolaus Copernicus probably consumed a lot of knowledge on how the Earth is the center of the universe etc. and probably nothing contradicting that notion. Can a LLM do that ?

And how many people actively do that? It's very rare we experience brilliance and often we stumble upon it by accident. Irrational behavior, coincidence or perhaps they were dropped on their heads when they were young.

Re: Notes on OpenAI's new o1 chain-of-thought models

#405
post #307
post #131

Earlier quoted context omitted.

> probabilistic token generators aren’t intelligence Maybe this has been extensively discussed before, but since I've lived under a rock: which parts of intelligence do you think are not representable as conditional probability distributions?

> which parts of intelligence do you think are not representable as conditional probability distributions Maybe I'm wrong here but a lot of our brilliance comes from acting against the statistical consensus. What I mean is, Nicolaus Copernicus probably consumed a lot of knowledge on how the Earth is the center of the universe etc. and probably nothing contradicting that notion. Can a LLM do that ?

Copernicus was an exception, not the rule. Would you say everyone else who lived at the time was not 'really' intelligent?

Re: Notes on OpenAI's new o1 chain-of-thought models

#406

I remember Murati's interview where she said about this PhD level reasoning and so on, so I was excited to see what they come up with - and it looks like they just used a bunch of models (like 4o's) and linked them in a chain of thought - which is exactly what we have been doing ourselves for a long time to get better results. So you have the usual disadvantages (time and money) and lose the only advantage you had wh…

do you know if someone actually compared a 4o CoT to the o1? I'm trying to find something on it, but I can't find anything.

LE: I found this tweet by Catena Labs of their MoA mix compared to o1-preview: https://x.com/catena_labs/status/1834416060071571836

Re: Notes on OpenAI's new o1 chain-of-thought models

#407
post #385

Earlier quoted context omitted.

1000% this. LLMs can't say "I don't know" because they don't actually think. I can coach a junior to get better. LLMs will just act like they know what they are doing and give the wrong results to people who aren't practitioners. Good on OAI calling their model Strawberry because of Internet trolls. Reactive vs proactive.

I get a lot of value out of ChatGPT but I also, fairly frequently, run into issues here. The real danger zones are areas that lie at or just beyond the edges of my own knowledge in a particular area. I'd say that most of my work use of ChatGPT does in fact save me time but, every so often, ChatGPT can still bullshit convincingly enough to waste an hour or two for me. The balance is still in its favour, but you have t…

Agreed, but the problem is if these things replace practitioners (what every MBA wants them to do), it's going to wreck the industry. Or maybe we'll get paid $$$$ to fix the problems they cause. GPT-4 introduced me to window functions in SQL (haven't written raw SQL in over a decade). But I'm experienced enough to look at window functions and compare them to subqueries and run some tests through the query planner to see what happens. That's knowledge that needs to be shared with the next generation of developers. And LLMs can't do that accurately.

Re: Notes on OpenAI's new o1 chain-of-thought models

#408
post #385

Earlier quoted context omitted.

1000% this. LLMs can't say "I don't know" because they don't actually think. I can coach a junior to get better. LLMs will just act like they know what they are doing and give the wrong results to people who aren't practitioners. Good on OAI calling their model Strawberry because of Internet trolls. Reactive vs proactive.

I get a lot of value out of ChatGPT but I also, fairly frequently, run into issues here. The real danger zones are areas that lie at or just beyond the edges of my own knowledge in a particular area. I'd say that most of my work use of ChatGPT does in fact save me time but, every so often, ChatGPT can still bullshit convincingly enough to waste an hour or two for me. The balance is still in its favour, but you have t…

This is basically the problem with all AI. It's good to a point, but they don't sufficiently know their limits/bounds and they will sometimes produce very odd results when you are right at those bounds.

AI in general just needs a way to identify when they're about to "make a coin flip" on an answer. With humans, we can quickly preference our asstalk with a disclaimer, at least.

Re: Notes on OpenAI's new o1 chain-of-thought models

#409
post #375
post #368

Earlier quoted context omitted.

I’d say the biggest difference between LLMs and SVMs is that a lot of people find LLMs useful on a daily basis. I’ve been using them almost daily for over two years now, and I keep on finding new things they can do that are useful to me.

Is there a post on your blog that lists your different uses of LLMs?

Not in a single place, but it came up in a podcast episode the other day - about 32 minutes in to this one I think https://softwaremisadventures.com/p/simon-willison-llm-weird...

Re: Notes on OpenAI's new o1 chain-of-thought models

#410
post #160

It’s still just a tool. It does not reason. It has some add-on logic the simulates it. We’re no closer to “AI” today than we were 20 years ago.

Personally I think “add-on logic that simulates reasoning” is a pretty good match for the “artificial” part of “artificial intelligence”. I’ve been tryin out the alternative term “initiation intelligence” recently, mainly to work around the baggage that’s become attached to the term AI.

Imitation intelligence, not initiation intelligence.
Post reply on HN