Live data from Hacker News

Notes on OpenAI's new o1 chain-of-thought models

simonwillison.net

611–620 of 659 posts

Re: Notes on OpenAI's new o1 chain-of-thought models

#611
post #601

Earlier quoted context omitted.

Says who? At a fundamental level

At a fundamental level, brains don’t operate on floating point numbers encoded in bits. They have chemicals to facilitate electrochemical reactions which can affect how they respond to input. They don’t throw away all knowledge of what they just said. They change continuously, not just in fixed training loops. They don’t operate in turns. I could go on. Honestly the number of people who just heard “learning,” “neural…

Fundamentally and physically are two different things. A logic gate is a logic gate if it's in neurons or silicon. Are abacus and calculators solving different things? No.

You're proving my point, things like them changing continuously are exactly what I mean when I say the brain is more efficient. Where there's a will theres a way and our brains are evidence that it can be done.

Re: Notes on OpenAI's new o1 chain-of-thought models

#612
post #438

Earlier quoted context omitted.

> The failure is in how you're using it. People, for the most part, know what they know and don't know. I am not uncertain that the distance between the earth and the sun varies, but I'm certain that I don't know the distance from the earth to the sun, at least not with better precision than about a light week. This is going to have to be fixed somehow to progress past where we are now with LLMs. Maybe expecting an L…

Empirically, they have reduced hallucinations. Where do OpenAI / Anthropic claim that their models won't hallucinate?

One example:

https://www.theverge.com/2024/3/28/24114664/microsoft-safety...

> Three features: Prompt Shields, which blocks prompt injections or malicious prompts from external documents that instruct models to go against their training; Groundedness Detection, which finds and blocks hallucinations; and safety evaluations, which assess model vulnerabilities, are now available in preview on Azure AI.

Re: Notes on OpenAI's new o1 chain-of-thought models

#613
post #442

Earlier quoted context omitted.

you're both on the wrong wavelength. No one has claimed it is better than an expert human yet. Be glad, for now your jobs are safe, why not use it as a tool to boost your productivity, yes, even though you'll get proportionally less use than others in other perhaps less "expert" jobs.

In order for it to boost productivity it needs to answer more than the regular questions for the top-3 languages on Stackoverflow, no? It often fails even for those questions. If I need to babysit it for every line of code, it's not a productivity boost.

Why does it need to answer more than that?

You underestimate the opportunity that exists for automation out there.

In my own case I've used it to make simple custom browser extensions transcribing PDFs, I don't have the time and wouldn't of made the effort to make the extension myself, the task would of continued to be done manually. It took two hours to make and it works, that's all I need in this case.

Perfection is the enemy of good.

Re: Notes on OpenAI's new o1 chain-of-thought models

#614
post #567

Earlier quoted context omitted.

Thank god we can finally end the scourge of interns to give the shareholders a little extra value. Good thing none of us ever started out as an intern.

I never said any of this will be good for society... In fact, I'm confident the current trajectory is going to cause wealth inequality at an entirely new level. Underestimating the impact these models can have is a risk I'm trying to expose...

I figured you weren't personally against interns.

More like, the prevailing attitude will be using AI to reduce labor costs at the lowest level, effectively gutting the ability to build a knowledge base for profit.

My snark was to add to that exposure.

Re: Notes on OpenAI's new o1 chain-of-thought models

#615

Earlier quoted context omitted.

The original riddle is of course: "A father and his son are in a car accident [...] When the boy is in hospital, the surgeon says: This is my child, I cannot operate on him". In the original riddle the answer is that the surgeon is female and the boy's mother. The riddle was supposed to point out gender stereotypes. So, as usual, ChatGPT fails to answer the modified riddle and gives the plagiarized stock answer and e…

> So, as usual, ChatGPT fails to answer the modified riddle and gives the plagiarized stock answer and explanation to the original one. No intelligence here. Or, fails in the same way any human would, when giving a snap answer to a riddle told to them on the fly - typically, a person would recognize a familiar riddle half of the first sentence in, and stop listening carefully, not expecting the other party to give th…

I'm curious what you think is happening here as your answer seems to imply it is thinking (and indeed rushing to an answer somehow). Do you think the generative AI has agency or a thought process? It doesn't seem to have anything approaching that to me, nor does it answer quickly.

It seems to be more like a weighing machine based on past tokens encountered together, so this is exactly the kind of answer we'd expect on a trivial question (I had no confusion over this question, my only confusion was why it was so basic).

It is surprisingly good at deceiving people and looking like it is thinking, when it only performs one of the many processes we use to think - pattern matching.

Re: Notes on OpenAI's new o1 chain-of-thought models

#616
post #438

Earlier quoted context omitted.

Empirically, they have reduced hallucinations. Where do OpenAI / Anthropic claim that their models won't hallucinate?

One example: https://www.theverge.com/2024/3/28/24114664/microsoft-safety... > Three features: Prompt Shields, which blocks prompt injections or malicious prompts from external documents that instruct models to go against their training; Groundedness Detection, which finds and blocks hallucinations; and safety evaluations, which assess model vulnerabilities, are now available in preview on Azure AI.

That wasn’t OpenAI making those claims, it was Microsoft Azure.

Re: Notes on OpenAI's new o1 chain-of-thought models

#617

Earlier quoted context omitted.

Phrased as it is, it deliberately gives away the answer by using the pronoun "he" for the doctor. The original deliberately obfuscates it by avoiding pronouns. So it doesn't take an understanding of gender roles, just grammar.

My point isn't that the model falls for gender stereotypes, but that it falls for thinking that it needs to solve the unmodified riddle. Humans fail at the original because they expect doctors to be male and miss crucial information because of that assumption. The model fails at the modification because it assumes that it is the unmodified riddle and misses crucial information because of that assumption. In both case…

They don't understand basic math or basic logic, so I don't think they understand grammar either.

They do understand/know the most likely words to follow on from a given word, which makes them very good at constructing convincing, plausible sentences in a given language - those sentences may well be gibberish or provably incorrect though - usually not because again most sentences in the dataset make some sort of sense, but sometimes the facade slips and it is apparent the GAI has no understanding and no theory of mind or even a basic model of relations between concepts (mother/father/son).

It is actually remarkable how like human writing their output is given how it is done, but there is no model of the world which backs their generated text which is a fatal flaw - as this example demonstrates.

Re: Notes on OpenAI's new o1 chain-of-thought models

#618

Earlier quoted context omitted.

Yes, this only helps multi-step reasoning. The model still has problems with general knowledge and deep facts. There's no way you can "reason" a correct answer to "list the tracklisting of some obscure 1991 demo by a band not on Wikipedia." You either know or you don't. I usually test new models with questions like "what are the levels in [semi-famous PC game from the 90s]?" The release version of GPT-4 could get abo…

It's actually much worse than that and you're inadvertently down playing how bad it is. It doesn't even know mildly obsecure facts that are on the internet. For example last night I was trying to do something with C# generics and it confidently told me I could use pattern matching on the type in a switch statwmnt, and threw out some convincing looking code. You can't, it's impossible. It wàa completely wrong. When I…

>>>As for non-proframming, we're about to see the birth of a new SEO movement of tricking AI models to believe your 'facts'.

This is kinda crazy to think about.

Re: Notes on OpenAI's new o1 chain-of-thought models

#619

Earlier quoted context omitted.

It's actually much worse than that and you're inadvertently down playing how bad it is. It doesn't even know mildly obsecure facts that are on the internet. For example last night I was trying to do something with C# generics and it confidently told me I could use pattern matching on the type in a switch statwmnt, and threw out some convincing looking code. You can't, it's impossible. It wàa completely wrong. When I…

>>>As for non-proframming, we're about to see the birth of a new SEO movement of tricking AI models to believe your 'facts'. This is kinda crazy to think about.

If you ask Google Gemini right now for the name of the whale in half moon bay harbor it will tell you it’s called Teresa T.

That was thanks to my experiment in influencing AI search: https://simonwillison.net/2024/Sep/8/teresa-t-whale-pillar-p...

Re: Notes on OpenAI's new o1 chain-of-thought models

#620

Earlier quoted context omitted.

> Treat it as a naive but intelligent intern That’s the problem: it’s a _terrible_ intern. A good intern will ask clarifying questions, tell me “I don’t know” or “I’m not sure I did it right”. LLMs do none of that, they will take whatever you ask and give a reasonable-sounding output that might be anything between brilliant and nonsense. With an intern, I don’t need to measure how good my prompting is, we’ll usually…

Makes me wonder if "I don't know" could be added to LLM: whenever an activation has no clear winner value (layman here), couldn't this indicate low response quality?

This exists and does work to some degree, e.g. Detecting hallucinations in large language models using semantic entropy https://www.nature.com/articles/s41586-024-07421-0
Post reply on HN