Live data from Hacker News

GPT-5.2

openai.com

831–840 of 1001 posts

Re: GPT-5.2

#831

Earlier quoted context omitted.

No-one should have the expectation LLMs are giving correct answers 100% of the time. It's inherent to the tech for them to be confidently wrong Code needs to be checked References need to be checked Any facts or claims need to be checked

According to the benchmarks here they're claiming up to 97% accuracy. That ought to be good enough to trust them right? Or maybe these benchmarks are all wrong

Something that is 97% accurate is wrong 3% of the time, so pointing out that it has gotten something wrong does not contradict 97% accuracy in the slightest.

Re: GPT-5.2

#832
post #560

Earlier quoted context omitted.

Why no grok 4.1 reasoning?

Do people other than Elon fans use grok? Honest question. I've never tried it.

I'm using Gemini in general, but Grok too. That's because sometimes Gemini Thinking is too slow, but Fast can get confused a lot. Grok strikes a nice balance between being quite smart (not Gemini 3 Pro level, but close) and very fast.

Re: GPT-5.2

#834
post #789

Earlier quoted context omitted.

Isn't that what no LLM can provide: being free of hallucinations?

I think the better word is confabulation; fabricating plausible but false narratives based on wrong memory. Fundamentally, these models try to produce plausible text. With language models getting large, they start creating internal world models, and some research shows they actually have truth dimensions. [0] I'm not an expert on the topic, but to me it sounds plausible that a good part of the problem of confabulatio…

> false narratives based on wrong memory

I don't think "wrong memory" is accurate, it's missing information and doesn't know it or is trained not to admit it.

Checkout the Dwarkesh Podcast episode https://www.dwarkesh.com/p/sholto-trenton-2 starting at 1:45:38

Here is the relevant quote by Trenton Bricken from the transcript:

One example I didn't talk about before with how the model retrieves facts: So you say, "What sport did Michael Jordan play?" And not only can you see it hop from like Michael Jordan to basketball and answer basketball. But the model also has an awareness of when it doesn't know the answer to a fact. And so, by default, it will actually say, "I don't know the answer to this question." But if it sees something that it does know the answer to, it will inhibit the "I don't know" circuit and then reply with the circuit that it actually has the answer to. So, for example, if you ask it, "Who is Michael Batkin?" —which is just a made-up fictional person— it will by default just say, "I don't know." It's only with Michael Jordan or someone else that it will then inhibit the "I don't know" circuit.

But what's really interesting here and where you can start making downstream predictions or reasoning about the model, is that the "I don't know" circuit is only on the name of the person. And so, in the paper we also ask it, "What paper did Andrej Karpathy write?" And so it recognizes the name Andrej Karpathy, because he's sufficiently famous, so that turns off the "I don't know" reply. But then when it comes time for the model to say what paper it worked on, it doesn't actually know any of his papers, and so then it needs to make something up. And so you can see different components and different circuits all interacting at the same time to lead to this final answer.

Re: GPT-5.2

#835

Earlier quoted context omitted.

For the record, brains are also not free of hallucinations.

That’s not a very useful observation though is it? The purpose of mechanisation is to standardise and over the long term reduce errors to zero. Otoh “The final truth is there is no truth”

A lot of mechanisation, especially in the modern world, is not deterministic and is not always 100% right; it's a fundamental "physics at scale" issue, not something new to LLMs. I think what happened when they first appeared was that people immediately clung to a superintelligence-type AI idea of what LLMs were supposed to do, then realised that's not what they are, then kept going and swung all the way over to "these things aren't good at anything really" or "if they only fix this ONE issue I have with them, they'll actually be useful"

Re: GPT-5.2

#836
post #769

A new model doesn't address the fundamental reliability issues with OpenAI's enterprise tier. As an enterprise customer, the experience has been disappointing. The platform is unstable, support is slow to respond even when escalated to account managers, and the UI is painfully slow to use. There are also baffling feature gaps, like the lack of connectors for custom GPTs. None of the major providers have a perfect ent…

Which tier are you? We are on the highest enterprise tier and I've found that OpenAI is a much more stable platform for high-usage than other providers. Can't say much about the UI though since I almost exclusively work with the API. I feel like UIs generally suck everywhere unless you want to do really generic stuff.

Re: GPT-5.2

#837

Earlier quoted context omitted.

> It's still a big issue that the models will make up plausible sounding but wrong or misleading explanations for things, and verifying their claims ends up taking time. And if it's a topic you don't care about enough, you might just end up misinformed. Exactly! One important thing LLMs have made me realise deeply is "No information" is better than false information. The way LLMs pull out completely incorrect explana…

> wrong or misleading explanations Exactly the same issue occurs with search. Unfortunately not everybody knows to mistrust AI responses, or have the skills to double-check information.

What is it about people making up lies to defend LLMs? In what world is it exactly the same as search? They're literally different things, since you get information from multiple sources and can do your own filtering.

Re: GPT-5.2

#838
post #828

Earlier quoted context omitted.

> It's still a big issue that the models will make up plausible sounding but wrong or misleading explanations for things, and verifying their claims ends up taking time. And if it's a topic you don't care about enough, you might just end up misinformed. Exactly! One important thing LLMs have made me realise deeply is "No information" is better than false information. The way LLMs pull out completely incorrect explana…

But most benchmarks are not about that... Are there even any "hallucination" public benchmarks?

"Benchmarks" for LLMs are a total hoax, since you can train them on the benchmarks themselves.

Re: GPT-5.2

#839
post #789

Earlier quoted context omitted.

Isn't that what no LLM can provide: being free of hallucinations?

I think the better word is confabulation; fabricating plausible but false narratives based on wrong memory. Fundamentally, these models try to produce plausible text. With language models getting large, they start creating internal world models, and some research shows they actually have truth dimensions. [0] I'm not an expert on the topic, but to me it sounds plausible that a good part of the problem of confabulatio…

No, the correct word is hallucinating. That's the word everyone uses and has been using. While it might not be technically correct, everyone knows what it means and more importantly, it's not a $3 word and everyone can relate to the concept. I also prefer all the _other_ more accurate alternative words Wikipedia offers to describe it:

"In the field of artificial intelligence (AI), a hallucination or artificial hallucination (also called bullshitting,[1][2] confabulation,[3] or delusion[4]) is"

Re: GPT-5.2

#840

Earlier quoted context omitted.

I added GPT-5.2 Pro to my pelican-alternatives benchmark for the first three prompts: Generate an SVG of an octopus operating a pipe organ Generate an SVG of a giraffe assembling a grandfather clock Generate an SVG of a starfish driving a bulldozer https://gally.net/temp/20251107pelican-alternatives/index.ht... GPT-5.2 Pro cost about 80 cents per prompt through OpenRouter, so I stopped there. I don’t feel like spendi…

Hi, it doesn't have Gemini 3.5 Pro which seems to be the best at this

That's probably because "Gemini 3.5 Pro" doesn't exist
Post reply on HN