Live data from Hacker News

GPT-5.5 hallucinates 3x more than MIT-licensed GLM-5.2

arrowtsx.dev

271–280 of 318 posts

Re: GPT-5.5 hallucinates 3x more than MIT-licensed GLM-5.2

#271
post #238

Earlier quoted context omitted.

https://x.com/polynoamial/status/2064210146558136827

I'd say that as OpenAI employee he's kinda biased on the topic

In what way is what he's saying wrong though? It makes him considerably more knowledgeable about the subject than the average Joe, and unfortunately for science, it's not like you're going to find an independent spherical researcher in the wild that exists without any bias.

Re: GPT-5.5 hallucinates 3x more than MIT-licensed GLM-5.2

#272

Earlier quoted context omitted.

I'd have to imagine there are wildly diminishing marginal returns to additional SFT/post-training passes. There are a bounded number of (useful) derivations/combinations of Duff's device. If Frontier Labs wish to reduce hallucinations on factual things, they will have to hire people (or the data providers will need to) to do fundamental research above and beyond what is available in extant literature and the web. IE…

As a side gig, I write novel software that solves problems no existing software does, that existing LLMs have difficulty reproducing, purely for the purpose of existing as LLM training data. There are journalists being hired to write Atlantic-worthy articles that exist only as LLM training data, because they're getting paid more than the Atlantic would pay them for it. It's insane. Yes, they are hiring the experts th…

> There are journalists being hired to write Atlantic-worthy articles that exist only as LLM training data, because they're getting paid more than the Atlantic would pay them for it.

I kinda doubt that quality-assertion of "Atlantic-worthy." While I have no doubt such articles are written solely as training data, I'd expect their quality to be much less than the real thing, since there's no public to critique them, probably little reputational risk for errors, and no professional ethics to uphold. Even if professional journalists were hired to do the work, I'd expect them to start phoning it in pretty quick, and skimp on fact-checking especially.

Re: GPT-5.5 hallucinates 3x more than MIT-licensed GLM-5.2

#273

Earlier quoted context omitted.

This would break reality. There’s some underlying physical law that prevents the existence of any algorithm of truth.

If you accept certain axioms a priori, it’s fine. If you simply let the machine intelligence take it for granted that induction works because nature is uniform and give it some way to test its predictions, it would have all the building blocks it needs to reason out a lot of very useful information. Which as the parent comment points out, people would absolutely pay a lot of money for.

If I am reading you correctly, that ends in testing, or essentially doing science.

That wouldn't satisfy the requirement to look at a claim and be able to tell if a claim was uncertainm roughly speaking.

Re: GPT-5.5 hallucinates 3x more than MIT-licensed GLM-5.2

#275
post #222

Earlier quoted context omitted.

>> it is clear that actual intelligence has plateaued significantly. > These are wild claims - Indeed, it is not clear there was any actual intelligence at any point. A lot of generated content sure, sometimes even useful, but not necessarily anything more.

What is the definition of "actual intelligence"? How does it differ from regular intelligence and non-intelligence? If someone can "design a custom asyncio event loop policy in that overrides get_child_watcher()", I would call that person intelligent. Does that mean that person is not actually intelligent but a mere content creation machine? Traditionally if you can create content, this shows you're intelligent. Crea…

Listing things that couldn't be done by tools until now but since then can is entirely unrelated to intelligence. I don't know even understand why it's being mentioned. A ruler can measure infinitely better than I can, nobody ever thought of it as intelligent. You can keep on giving more (genuinely) impressive examples yet it's not helping at all.

Re: GPT-5.5 hallucinates 3x more than MIT-licensed GLM-5.2

#277

Earlier quoted context omitted.

As a side gig, I write novel software that solves problems no existing software does, that existing LLMs have difficulty reproducing, purely for the purpose of existing as LLM training data. There are journalists being hired to write Atlantic-worthy articles that exist only as LLM training data, because they're getting paid more than the Atlantic would pay them for it. It's insane. Yes, they are hiring the experts th…

> There are journalists being hired to write Atlantic-worthy articles that exist only as LLM training data, because they're getting paid more than the Atlantic would pay them for it. I kinda doubt that quality-assertion of "Atlantic-worthy." While I have no doubt such articles are written solely as training data, I'd expect their quality to be much less than the real thing, since there's no public to critique them, p…

That's where the review process comes in.

I've found the review process for these things to be far more vicious and demanding than in the real world.

Re: GPT-5.5 hallucinates 3x more than MIT-licensed GLM-5.2

#279

Earlier quoted context omitted.

I'd have to imagine there are wildly diminishing marginal returns to additional SFT/post-training passes. There are a bounded number of (useful) derivations/combinations of Duff's device. If Frontier Labs wish to reduce hallucinations on factual things, they will have to hire people (or the data providers will need to) to do fundamental research above and beyond what is available in extant literature and the web. IE…

As a side gig, I write novel software that solves problems no existing software does, that existing LLMs have difficulty reproducing, purely for the purpose of existing as LLM training data. There are journalists being hired to write Atlantic-worthy articles that exist only as LLM training data, because they're getting paid more than the Atlantic would pay them for it. It's insane. Yes, they are hiring the experts th…

>There are journalists being hired to write Atlantic-worthy articles that exist only as LLM training data, because they're getting paid more than the Atlantic would pay them for it.

So now part of the training is directly slop - just of the pre-AI variety?

Re: GPT-5.5 hallucinates 3x more than MIT-licensed GLM-5.2

#280

Earlier quoted context omitted.

Not the worst way to make money, but if internet-scale data were not enough to reduce errors to a somewhat tolerable margin, how much data do they hope to collect in this manner?

Right now, this is a 10-figure run rate industry. They are generating a lot of this. Also remember it's not just quantity, it's roughly active learning - they're paying for training data that's at the classification boundary, which is way more valuable. I have gotten offers for contracts for full time jobs at high rates with AI labs to do this. Meta has reallocated a lot of their full time SWE staff to do this. All o…

Sound like the very definition of marginal returns and/or desperation, combined with throwing money at the problem...
Post reply on HN