Earlier quoted context omitted.
No-one should have the expectation LLMs are giving correct answers 100% of the time. It's inherent to the tech for them to be confidently wrong Code needs to be checked References need to be checked Any facts or claims need to be checked
According to the benchmarks here they're claiming up to 97% accuracy. That ought to be good enough to trust them right? Or maybe these benchmarks are all wrong
GPT-5.2
831–840 of 1001 posts
Re: GPT-5.2
#832Earlier quoted context omitted.
Why no grok 4.1 reasoning?
Do people other than Elon fans use grok? Honest question. I've never tried it.
Re: GPT-5.2
#833Re: GPT-5.2
#834Earlier quoted context omitted.
Isn't that what no LLM can provide: being free of hallucinations?
I think the better word is confabulation; fabricating plausible but false narratives based on wrong memory. Fundamentally, these models try to produce plausible text. With language models getting large, they start creating internal world models, and some research shows they actually have truth dimensions. [0] I'm not an expert on the topic, but to me it sounds plausible that a good part of the problem of confabulatio…
I don't think "wrong memory" is accurate, it's missing information and doesn't know it or is trained not to admit it.
Checkout the Dwarkesh Podcast episode https://www.dwarkesh.com/p/sholto-trenton-2 starting at 1:45:38
Here is the relevant quote by Trenton Bricken from the transcript:
One example I didn't talk about before with how the model retrieves facts: So you say, "What sport did Michael Jordan play?" And not only can you see it hop from like Michael Jordan to basketball and answer basketball. But the model also has an awareness of when it doesn't know the answer to a fact. And so, by default, it will actually say, "I don't know the answer to this question." But if it sees something that it does know the answer to, it will inhibit the "I don't know" circuit and then reply with the circuit that it actually has the answer to. So, for example, if you ask it, "Who is Michael Batkin?" —which is just a made-up fictional person— it will by default just say, "I don't know." It's only with Michael Jordan or someone else that it will then inhibit the "I don't know" circuit.
But what's really interesting here and where you can start making downstream predictions or reasoning about the model, is that the "I don't know" circuit is only on the name of the person. And so, in the paper we also ask it, "What paper did Andrej Karpathy write?" And so it recognizes the name Andrej Karpathy, because he's sufficiently famous, so that turns off the "I don't know" reply. But then when it comes time for the model to say what paper it worked on, it doesn't actually know any of his papers, and so then it needs to make something up. And so you can see different components and different circuits all interacting at the same time to lead to this final answer.
Re: GPT-5.2
#835Earlier quoted context omitted.
For the record, brains are also not free of hallucinations.
That’s not a very useful observation though is it? The purpose of mechanisation is to standardise and over the long term reduce errors to zero. Otoh “The final truth is there is no truth”
Re: GPT-5.2
#836A new model doesn't address the fundamental reliability issues with OpenAI's enterprise tier. As an enterprise customer, the experience has been disappointing. The platform is unstable, support is slow to respond even when escalated to account managers, and the UI is painfully slow to use. There are also baffling feature gaps, like the lack of connectors for custom GPTs. None of the major providers have a perfect ent…
Re: GPT-5.2
#837Earlier quoted context omitted.
> It's still a big issue that the models will make up plausible sounding but wrong or misleading explanations for things, and verifying their claims ends up taking time. And if it's a topic you don't care about enough, you might just end up misinformed. Exactly! One important thing LLMs have made me realise deeply is "No information" is better than false information. The way LLMs pull out completely incorrect explana…
> wrong or misleading explanations Exactly the same issue occurs with search. Unfortunately not everybody knows to mistrust AI responses, or have the skills to double-check information.
Re: GPT-5.2
#838Earlier quoted context omitted.
> It's still a big issue that the models will make up plausible sounding but wrong or misleading explanations for things, and verifying their claims ends up taking time. And if it's a topic you don't care about enough, you might just end up misinformed. Exactly! One important thing LLMs have made me realise deeply is "No information" is better than false information. The way LLMs pull out completely incorrect explana…
But most benchmarks are not about that... Are there even any "hallucination" public benchmarks?
Re: GPT-5.2
#839Earlier quoted context omitted.
Isn't that what no LLM can provide: being free of hallucinations?
I think the better word is confabulation; fabricating plausible but false narratives based on wrong memory. Fundamentally, these models try to produce plausible text. With language models getting large, they start creating internal world models, and some research shows they actually have truth dimensions. [0] I'm not an expert on the topic, but to me it sounds plausible that a good part of the problem of confabulatio…
"In the field of artificial intelligence (AI), a hallucination or artificial hallucination (also called bullshitting,[1][2] confabulation,[3] or delusion[4]) is"
Re: GPT-5.2
#840Earlier quoted context omitted.
I added GPT-5.2 Pro to my pelican-alternatives benchmark for the first three prompts: Generate an SVG of an octopus operating a pipe organ Generate an SVG of a giraffe assembling a grandfather clock Generate an SVG of a starfish driving a bulldozer https://gally.net/temp/20251107pelican-alternatives/index.ht... GPT-5.2 Pro cost about 80 cents per prompt through OpenRouter, so I stopped there. I don’t feel like spendi…
Hi, it doesn't have Gemini 3.5 Pro which seems to be the best at this