Earlier quoted context omitted.
I ask for confidence scores in my custom instructions / prompts, and LLMs do surprisingly well at estimating their own knowledge most of the time.
I’m with the people pushing back on the “confidence scores” framing, but I think the deeper issue is that we’re still stuck in the wrong mental model. It’s tempting to think of a language model as a shallow search engine that happens to output text, but that metaphor doesn’t actually match what’s happening under the hood. A model doesn’t “know” facts or measure uncertainty in a Bayesian sense. All it really does is t…
GPT-5.2
941–950 of 1001 posts
Re: GPT-5.2
#942In my experience, the best models are already nearly as good as you can be for a large fraction of what I personally use them for, which is basically as a more efficient search engine. The thing that would now make the biggest difference isn't "more intelligence", whatever that might mean, but better grounding. It's still a big issue that the models will make up plausible sounding but wrong or misleading explanations…
Re: GPT-5.2
#943Earlier quoted context omitted.
Isn't that what no LLM can provide: being free of hallucinations?
For the record, brains are also not free of hallucinations.
LLMs are being sold as viable replacement of paid employees.
If they were not, they wouldn’t be funded the way they are.
Re: GPT-5.2
#944Somewhat tangential: The second link says "System card": https://cdn.openai.com/pdf/3a4153c8-c748-4b71-8e31-aecbde944... Does that term have special meaning in the AI/LLM world? I never heard it before. I Google'd the term "System Card LLM" and got a bunch of hits. I am so surprised that I never saw the term used here in HN before. Also, the layout looks exactly like a scientific paper written in LaTeX. Who is the ex…
Re: GPT-5.2
#945Earlier quoted context omitted.
I’m with the people pushing back on the “confidence scores” framing, but I think the deeper issue is that we’re still stuck in the wrong mental model. It’s tempting to think of a language model as a shallow search engine that happens to output text, but that metaphor doesn’t actually match what’s happening under the hood. A model doesn’t “know” facts or measure uncertainty in a Bayesian sense. All it really does is t…
You have a subtle slight of hand. You use the word “plausible” instead of “correct.”
Re: GPT-5.2
#946Earlier quoted context omitted.
Most non tech people I talked with don't care at all about LLMs. They also are not impressed at all ("Okay, that's like google and internet").
Old people? I think it would be hard to find a lot of people under 20 who don't use ChatGPT daily. At least among ones that are still studying.
I have hard time to imagine why non-tech people would find a use for LLMs, let's say nothing in your life forces you to produce information (be it textual, pictural or anything that can be related to information). Let's say your needs are focused on spending good times with friends or your family, eating nice dishes (home cooked or restaurant), spending your money on furnitures, rents, clothes, tools and etc.
Why would you need an AI that produce information in an information-bloated world ?
You probably met someone that "fell in love with woodworking" or idk, after having watched youtube videos (that person probably built a chair, a table or something akin). I don't think stuff like "Hi, I have these materials, what can I do with it" produce more interesting results than just nerding on the internet or in a library looking for references (on japaneese handcrafted furnitures, vintage ikea designs, old school woodworking, ...). (Or maybe the LLM will be able to give you a list of good reads, which is nice but somewhat of a limited and basic use).
Agentic AI and more efficient/intelligent AIs are not very interesting for people like and are at best a proxy for otherly findable information. Of course, not everyone is like , the majority of people don't even need to invest time in a "creative" hobby and instead they will watch movies, invest time in sport, invest time in sociability, go to museums, read books; you could imagine having AIs that write books, invent films, invent artworks, talk with you, but I am pretty sure that there is something more than just "watch a movie" or "read a book" when performing these activities; as someone who likes reading or watching movies, what I enjoy is following the evolutions of the authors of the pieces, understanding their posture toward its ancestors, its era-mates, toward its own previous visions and whatnot. I enjoy to find a movie "weird" "goofy" "sublime" and whatnot, because I enjoy a small amount of parasociality with the authors and am finally brought to say things like "Ahah, Lynch was such a weirdo when he shot Blue Velvet" (okay, maybe not that type of bully judgement, but you may be understanding what I mean).
I think I would find it uninspiring to read an AI written book, because I couldn't live this small parasocial experience. Maybe you could get me with music, but I still think there's a lot of activity in loving a song. I love Bach, but am pretty sure also I like Bach the character (from what I speculate from the songs I listen). I imagine that guy in front of his keyboard, having the chance to live a -weird- moment of extasy when he produces the best lines of the chaconne (if he was living in our times he would relisten to what he produced again and again and nodding to himself "man, that's sick").
What could I experience from an LLM ? "Here is the perfect novel I wrote specifically for you based on your tastes:". There would be no imaginary Bach that I would like to drink a beer with, no testimony of a human reaching the state of mind in which you produce an absolute (in fact highly relative, but you need to lie to yourself) "hit".
All of this is highly personnal, but I would be curious to know what others think.
Re: GPT-5.2
#947A year ago Sunday Pichai declared code red, now it’s Sam Altman declaring code red. How tables have turned, and I think the acquisition of Windsurf and Kevin Hou by Google seems to correlate with their level up.
To make an argument it was Kevin Hou, then we would need to see Antigravity their new IDE being key. I think the crown jewel are the Gemini models.
Re: GPT-5.2
#948Earlier quoted context omitted.
I’m with the people pushing back on the “confidence scores” framing, but I think the deeper issue is that we’re still stuck in the wrong mental model. It’s tempting to think of a language model as a shallow search engine that happens to output text, but that metaphor doesn’t actually match what’s happening under the hood. A model doesn’t “know” facts or measure uncertainty in a Bayesian sense. All it really does is t…
Hallucinations are a feature of reality that LLMs have inherited. It’s amazing that experts like yourself who have a good grasp of the manifold MoE configuration don’t get that. LLMs much like humans weight high dimensionality across the entire model then manifold then string together an attentive answer best weighted. Just like your doctor occasionally giving you wrong advice too quickly so does this sometimes eithe…
Really? When I search for cases on LexisNexis, it does not return made-up cases which do not actually exist.
Re: GPT-5.2
#949Earlier quoted context omitted.
I’m with the people pushing back on the “confidence scores” framing, but I think the deeper issue is that we’re still stuck in the wrong mental model. It’s tempting to think of a language model as a shallow search engine that happens to output text, but that metaphor doesn’t actually match what’s happening under the hood. A model doesn’t “know” facts or measure uncertainty in a Bayesian sense. All it really does is t…
Hallucinations are a feature of reality that LLMs have inherited. It’s amazing that experts like yourself who have a good grasp of the manifold MoE configuration don’t get that. LLMs much like humans weight high dimensionality across the entire model then manifold then string together an attentive answer best weighted. Just like your doctor occasionally giving you wrong advice too quickly so does this sometimes eithe…
Huh? Are you arguing that we still live in a pre-scientific era where there’s no way to measure truth?
As a simple example, I asked Google about houseplant biology recently. The answer was very confidently wrong telling me that spider plants have a particular metabolic pathway because it confused them with jade plants and the two are often mentioned together. Humans wouldn’t make this mistake because they’d either know the answer or say that they don’t. LLMs do that constantly because they lack understanding and metacognitive abilities.
Re: GPT-5.2
#950Earlier quoted context omitted.
Is Adaptive Reasoning gone from GPT-5.2? It was a big part of the release of 5.1 and Codex-Max. Really felt like the future.
Yes, GPT-5.2 still has adaptive reasoning - we just didn't call it out by name this time. Like 5.1 and codex-max, it should do a better job at answering quickly on easy queries and taking its time on harder queries.
Extended and heavy are about raising the floor (~25% and ~45% or some other ratio respectively) not determining the ceiling.