Live data from Hacker News

GPT-5.2

openai.com

941–950 of 1001 posts

Re: GPT-5.2

#941
post #830

Earlier quoted context omitted.

I ask for confidence scores in my custom instructions / prompts, and LLMs do surprisingly well at estimating their own knowledge most of the time.

I’m with the people pushing back on the “confidence scores” framing, but I think the deeper issue is that we’re still stuck in the wrong mental model. It’s tempting to think of a language model as a shallow search engine that happens to output text, but that metaphor doesn’t actually match what’s happening under the hood. A model doesn’t “know” facts or measure uncertainty in a Bayesian sense. All it really does is t…

It's not even a manifold https://arxiv.org/abs/2504.01002

Re: GPT-5.2

#942
post #765

In my experience, the best models are already nearly as good as you can be for a large fraction of what I personally use them for, which is basically as a more efficient search engine. The thing that would now make the biggest difference isn't "more intelligence", whatever that might mean, but better grounding. It's still a big issue that the models will make up plausible sounding but wrong or misleading explanations…

All of them are heavily invested in improving grounding. The money isn't on personal use but enterprise customers and for those, grounding is essential.

Re: GPT-5.2

#943

Earlier quoted context omitted.

Isn't that what no LLM can provide: being free of hallucinations?

For the record, brains are also not free of hallucinations.

How much do you hallucinate at work? How many of your work hallucinations do you confidently present as reality in communication or code?

LLMs are being sold as viable replacement of paid employees.

If they were not, they wouldn’t be funded the way they are.

Re: GPT-5.2

#944

Somewhat tangential: The second link says "System card": https://cdn.openai.com/pdf/3a4153c8-c748-4b71-8e31-aecbde944... Does that term have special meaning in the AI/LLM world? I never heard it before. I Google'd the term "System Card LLM" and got a bunch of hits. I am so surprised that I never saw the term used here in HN before. Also, the layout looks exactly like a scientific paper written in LaTeX. Who is the ex…

Yeah, search HN for the term. It's a relatively big topic of conversation.

Re: GPT-5.2

#945

Earlier quoted context omitted.

I’m with the people pushing back on the “confidence scores” framing, but I think the deeper issue is that we’re still stuck in the wrong mental model. It’s tempting to think of a language model as a shallow search engine that happens to output text, but that metaphor doesn’t actually match what’s happening under the hood. A model doesn’t “know” facts or measure uncertainty in a Bayesian sense. All it really does is t…

You have a subtle slight of hand. You use the word “plausible” instead of “correct.”

Do you have a better word that describes "things that look correct without definitely being so"? I think "plausible" is the perfect word for that. It's not a sleight of hand to use a word that is exactly defined as the intention.

Re: GPT-5.2

#946
post #745

Earlier quoted context omitted.

Most non tech people I talked with don't care at all about LLMs. They also are not impressed at all ("Okay, that's like google and internet").

Old people? I think it would be hard to find a lot of people under 20 who don't use ChatGPT daily. At least among ones that are still studying.

I wanted to reflect a bit on this.

I have hard time to imagine why non-tech people would find a use for LLMs, let's say nothing in your life forces you to produce information (be it textual, pictural or anything that can be related to information). Let's say your needs are focused on spending good times with friends or your family, eating nice dishes (home cooked or restaurant), spending your money on furnitures, rents, clothes, tools and etc.

Why would you need an AI that produce information in an information-bloated world ?

You probably met someone that "fell in love with woodworking" or idk, after having watched youtube videos (that person probably built a chair, a table or something akin). I don't think stuff like "Hi, I have these materials, what can I do with it" produce more interesting results than just nerding on the internet or in a library looking for references (on japaneese handcrafted furnitures, vintage ikea designs, old school woodworking, ...). (Or maybe the LLM will be able to give you a list of good reads, which is nice but somewhat of a limited and basic use).

Agentic AI and more efficient/intelligent AIs are not very interesting for people like and are at best a proxy for otherly findable information. Of course, not everyone is like , the majority of people don't even need to invest time in a "creative" hobby and instead they will watch movies, invest time in sport, invest time in sociability, go to museums, read books; you could imagine having AIs that write books, invent films, invent artworks, talk with you, but I am pretty sure that there is something more than just "watch a movie" or "read a book" when performing these activities; as someone who likes reading or watching movies, what I enjoy is following the evolutions of the authors of the pieces, understanding their posture toward its ancestors, its era-mates, toward its own previous visions and whatnot. I enjoy to find a movie "weird" "goofy" "sublime" and whatnot, because I enjoy a small amount of parasociality with the authors and am finally brought to say things like "Ahah, Lynch was such a weirdo when he shot Blue Velvet" (okay, maybe not that type of bully judgement, but you may be understanding what I mean).

I think I would find it uninspiring to read an AI written book, because I couldn't live this small parasocial experience. Maybe you could get me with music, but I still think there's a lot of activity in loving a song. I love Bach, but am pretty sure also I like Bach the character (from what I speculate from the songs I listen). I imagine that guy in front of his keyboard, having the chance to live a -weird- moment of extasy when he produces the best lines of the chaconne (if he was living in our times he would relisten to what he produced again and again and nodding to himself "man, that's sick").

What could I experience from an LLM ? "Here is the perfect novel I wrote specifically for you based on your tastes:". There would be no imaginary Bach that I would like to drink a beer with, no testimony of a human reaching the state of mind in which you produce an absolute (in fact highly relative, but you need to lie to yourself) "hit".

All of this is highly personnal, but I would be curious to know what others think.

Re: GPT-5.2

#947
post #381

A year ago Sunday Pichai declared code red, now it’s Sam Altman declaring code red. How tables have turned, and I think the acquisition of Windsurf and Kevin Hou by Google seems to correlate with their level up.

Acquisition of noam shazeer to supercharge their Gemini flagship model line I think made a bigger impact.

To make an argument it was Kevin Hou, then we would need to see Antigravity their new IDE being key. I think the crown jewel are the Gemini models.

Re: GPT-5.2

#948

Earlier quoted context omitted.

I’m with the people pushing back on the “confidence scores” framing, but I think the deeper issue is that we’re still stuck in the wrong mental model. It’s tempting to think of a language model as a shallow search engine that happens to output text, but that metaphor doesn’t actually match what’s happening under the hood. A model doesn’t “know” facts or measure uncertainty in a Bayesian sense. All it really does is t…

Hallucinations are a feature of reality that LLMs have inherited. It’s amazing that experts like yourself who have a good grasp of the manifold MoE configuration don’t get that. LLMs much like humans weight high dimensionality across the entire model then manifold then string together an attentive answer best weighted. Just like your doctor occasionally giving you wrong advice too quickly so does this sometimes eithe…

> Hallucinations are a feature of reality that LLMs have inherited.

Really? When I search for cases on LexisNexis, it does not return made-up cases which do not actually exist.

Re: GPT-5.2

#949

Earlier quoted context omitted.

I’m with the people pushing back on the “confidence scores” framing, but I think the deeper issue is that we’re still stuck in the wrong mental model. It’s tempting to think of a language model as a shallow search engine that happens to output text, but that metaphor doesn’t actually match what’s happening under the hood. A model doesn’t “know” facts or measure uncertainty in a Bayesian sense. All it really does is t…

Hallucinations are a feature of reality that LLMs have inherited. It’s amazing that experts like yourself who have a good grasp of the manifold MoE configuration don’t get that. LLMs much like humans weight high dimensionality across the entire model then manifold then string together an attentive answer best weighted. Just like your doctor occasionally giving you wrong advice too quickly so does this sometimes eithe…

> Hallucinations are a feature of reality that LLMs have inherited.

Huh? Are you arguing that we still live in a pre-scientific era where there’s no way to measure truth?

As a simple example, I asked Google about houseplant biology recently. The answer was very confidently wrong telling me that spider plants have a particular metabolic pathway because it confused them with jade plants and the two are often mentioned together. Humans wouldn’t make this mistake because they’d either know the answer or say that they don’t. LLMs do that constantly because they lack understanding and metacognitive abilities.

Re: GPT-5.2

#950

Earlier quoted context omitted.

Is Adaptive Reasoning gone from GPT-5.2? It was a big part of the release of 5.1 and Codex-Max. Really felt like the future.

Yes, GPT-5.2 still has adaptive reasoning - we just didn't call it out by name this time. Like 5.1 and codex-max, it should do a better job at answering quickly on easy queries and taking its time on harder queries.

Why have "light" or "low" thinking then? I've mentioned this before in other places, but there should only be "none," "standard," "extended," and maybe "heavy."

Extended and heavy are about raising the floor (~25% and ~45% or some other ratio respectively) not determining the ceiling.

Post reply on HN