Live data from Hacker News

GPT-5.2

openai.com

811–820 of 1001 posts

Re: GPT-5.2

#811

Earlier quoted context omitted.

> wrong or misleading explanations Exactly the same issue occurs with search. Unfortunately not everybody knows to mistrust AI responses, or have the skills to double-check information.

No, it's not the same. Search results send/show you one or more specific pages/websites. And each website has a different trust factor. Yes, plenty of people repeat things they "read on the Internet" as truths, but it's easy to debunk some of them just based on the site reputation. With AI responses, the reputation is shared with the good answers as well, because they do give good answers most of the time, but also h…

Community notes on X seems to be one of the highest profile recent experiments trying to address this issue

Re: GPT-5.2

#812
post #560

Earlier quoted context omitted.

Why no grok 4.1 reasoning?

Do people other than Elon fans use grok? Honest question. I've never tried it.

Unlike openai, you can use the latest grok models without verifying your organization and giving your ID.

Re: GPT-5.2

#813
post #780

Earlier quoted context omitted.

Isn't that what no LLM can provide: being free of hallucinations?

Yes, they'll probably not go away, but it's got to be possible to handle them better. Gemini (the app) has a "mitigation" feature where it tries to to Google searches to support its statements. That doesn't currently work properly in my experience. It also seems to be doing something where it adds references to statements (With a separate model? With a second pass over the output? Not sure how that works.). That work…

[deleted]

Re: GPT-5.2

#814

Earlier quoted context omitted.

Yep, the point we wanted to make here is that GPT-5.2's vision is better, not perfect. Cherrypicking a perfect output would actually mislead readers, and that wasn't our intent.

That would be a laudable goal, but I feel like it's contradicted by the text: > Even on a low-quality image, GPT‑5.2 identifies the main regions and places boxes that roughly match the true locations of each component I would not consider it to have "identified the main regions" or to have "roughly matched the true locations" when ~1/3 of the boxes have incorrect labels . The remark "even on a low-quality image" is n…

Leave it to OpenAI to be dishonest about being dishonest. It seems they're also editing this post without notice as well.

Re: GPT-5.2

#815
post #125

Wow, there's a lot going on with this pelican riding a bicycle: https://gist.github.com/simonw/c31d7afc95fe6b40506a9562b5e83...

What is good at SVG design?

Ive not seen any model being good in graphic/svg creation so far - all of the stuff mostly looks ugly and somewhat "synthetic-disorted".

And lately, Claude (web) started to draw ascii charts from one day to another indstead of colorful infographicstyled-images as it did before (they were only slightly better than the ascii charts)

Re: GPT-5.2

#816
post #765

In my experience, the best models are already nearly as good as you can be for a large fraction of what I personally use them for, which is basically as a more efficient search engine. The thing that would now make the biggest difference isn't "more intelligence", whatever that might mean, but better grounding. It's still a big issue that the models will make up plausible sounding but wrong or misleading explanations…

I agree, but the question is how better grounding can be achieved without a major research breakthrough.

I believe the real issue is that LLMs are still so bad at reasoning. In my experience, the worst hallucinations occur where only handful of sources exist for some set of facts (e.g laws of small countries or descriptions of niche products).

LLMs know these sources and they refer to them but they are interpreting them incorrectly. They are incapable of focusing on the semantics of one specific page because they get "distracted" by their pattern matching nature.

Now people will say that this is unavoidable given the way in which transformers work. And this is true.

But shouldn't it be possible to include some measure of data sparsity in the training so that models know when they don't know enough? That would enable them to boost the weight of the context (including sources they find through inference time search/RAG) relative to to their pretraining.

Re: GPT-5.2

#817
post #780

Earlier quoted context omitted.

Isn't that what no LLM can provide: being free of hallucinations?

Yes, they'll probably not go away, but it's got to be possible to handle them better. Gemini (the app) has a "mitigation" feature where it tries to to Google searches to support its statements. That doesn't currently work properly in my experience. It also seems to be doing something where it adds references to statements (With a separate model? With a second pass over the output? Not sure how that works.). That work…

Doubt it. I suspect it’s fundamentally not possible in the spirit you intend it.

Reality is perfectly fine with deception and inaccuracy. For language to magically be self constraining enough to only make verified statements is… impossible.

Re: GPT-5.2

#818

Earlier quoted context omitted.

> It's still a big issue that the models will make up plausible sounding but wrong or misleading explanations for things, and verifying their claims ends up taking time. And if it's a topic you don't care about enough, you might just end up misinformed. Exactly! One important thing LLMs have made me realise deeply is "No information" is better than false information. The way LLMs pull out completely incorrect explana…

> wrong or misleading explanations Exactly the same issue occurs with search. Unfortunately not everybody knows to mistrust AI responses, or have the skills to double-check information.

[deleted]

Re: GPT-5.2

#819

Earlier quoted context omitted.

For the record, brains are also not free of hallucinations.

I still don’t really get this argument/excuse for why it’s acceptable that LLMs hallucinate. These tools are meant to support us, but we end up with two parties who are, as you say, prone to “hallucination” and it becomes a situation of the blind leading the blind. Ideally in these scenarios there’s at least one party with a definitive or deterministic view so the other party (i.e. us) at least has some trust in the…

[deleted]

Re: GPT-5.2

#820
post #426

Earlier quoted context omitted.

You are absolutely right!

Someone didn't think so, lol. I debated not saying anything because the AI partisans are just so awful.

I think the above comment was a joke (Claude frequently says that whenever you challenge it, whether you are right or wrong)
Post reply on HN