Live data from Hacker News

GPT-5.5 hallucinates 3x more than MIT-licensed GLM-5.2

arrowtsx.dev

261–270 of 318 posts

Re: GPT-5.5 hallucinates 3x more than MIT-licensed GLM-5.2

#261
post #96

Earlier quoted context omitted.

It’s not as simple. I trained an LLM before on exactly this, to scratch the itch of this question. The task was simple, using the MS-MARCO[0] dataset which contains queries, search results, answers, I made a training set that has: 1. Questions paired with real results supporting them (mixed with some irrelevant results), and a correct answer 2. Questions paired only with irrelevant results, with the answer “No answer…

Thank you for sharing! Based on your experience, do you think a two-model system might fare better? For example, two models in serial where the second model is trained to "sniff out" potential hallucinations and fact check them (and possibly iterate with the first model)?

I do think it might improve but only marginally.

You are however likely to observe better results in smaller models since they're usually more strapped for "cognitive capacity", so two separate calls reduce the load in each request, and hallucination in my experience is a common side effect of overloading an LLM cognitively.

Re: GPT-5.5 hallucinates 3x more than MIT-licensed GLM-5.2

#262
post #231

Earlier quoted context omitted.

Sucky human-written code is still based on human understanding, which can change over time, be readjusted or solidified. People implement something wrong once, then update their perspective, then in the future does it right. LLMs doesn't have this benefit. You forget to add the correct to the system prompt, and the LLM will repeat the same mistake over and over, and worse than that, their mistakes aren't based on the…

> while a human you can more or less dump them everything you sit on, and let them shift it through, and they'll mostly make it out OK i dont see why software engineers are paid so well, and are so hard to hire? just dump a bunch of requirements on a homeless person and itll just work out

I have no idea what point you're making here.

Re: GPT-5.5 hallucinates 3x more than MIT-licensed GLM-5.2

#263

Earlier quoted context omitted.

Right now, this is a 10-figure run rate industry. They are generating a lot of this. Also remember it's not just quantity, it's roughly active learning - they're paying for training data that's at the classification boundary, which is way more valuable. I have gotten offers for contracts for full time jobs at high rates with AI labs to do this. Meta has reallocated a lot of their full time SWE staff to do this. All o…

Yes, they do have money to burn, and this will bring some improvements for sure, but active learning has never really worked out, has it? And even 10% of the educated population doing this for, like, 50 years is not that much data, while normally each accuracy percentage is more and more data-expensive.

Eh, maybe 20% of my ML career earnings has been showing that active learning works out. shrug

Re: GPT-5.5 hallucinates 3x more than MIT-licensed GLM-5.2

#264

Earlier quoted context omitted.

As a side gig, I write novel software that solves problems no existing software does, that existing LLMs have difficulty reproducing, purely for the purpose of existing as LLM training data. There are journalists being hired to write Atlantic-worthy articles that exist only as LLM training data, because they're getting paid more than the Atlantic would pay them for it. It's insane. Yes, they are hiring the experts th…

What kind of programs? Can you give an example of the tasks?

I cannot give examples.

For the more interesting contracts, I create examples from whole cloth.

Re: GPT-5.5 hallucinates 3x more than MIT-licensed GLM-5.2

#265

Earlier quoted context omitted.

I find these internet arguments talking about LLMs as if they are trained by reading the internet to be wild. Yes, pretraining still exists. But for the past few years, pretraining by reading the internet is just the initial bootstrapping of LLM training. The RL training they get from bespoke training data, with very very different characteristics than what these armchair analyses claim, dominates these days.

I'd have to imagine there are wildly diminishing marginal returns to additional SFT/post-training passes. There are a bounded number of (useful) derivations/combinations of Duff's device. If Frontier Labs wish to reduce hallucinations on factual things, they will have to hire people (or the data providers will need to) to do fundamental research above and beyond what is available in extant literature and the web. IE…

>> They need expert reviewers in virtually every interesting topic, which fundamentally is an intractable problem, especially since things change all the time.

How odd. It's Expert Systems and the Knowledge Acquisition Bottleneck all over again.

Re: GPT-5.5 hallucinates 3x more than MIT-licensed GLM-5.2

#266

Earlier quoted context omitted.

Not the worst way to make money, but if internet-scale data were not enough to reduce errors to a somewhat tolerable margin, how much data do they hope to collect in this manner?

Right now, this is a 10-figure run rate industry. They are generating a lot of this. Also remember it's not just quantity, it's roughly active learning - they're paying for training data that's at the classification boundary, which is way more valuable. I have gotten offers for contracts for full time jobs at high rates with AI labs to do this. Meta has reallocated a lot of their full time SWE staff to do this. All o…

I think this all reinforces the idea that the industry has no idea how to pursue general intelligence. Hence vast sums being spent plugging holes and fitting the models to more and more specific tasks.

But with this approach, there will always be the next car wash test showing that it is an illusion. It seems to me the limits of the Bitter Lesson are showing.

Re: GPT-5.5 hallucinates 3x more than MIT-licensed GLM-5.2

#267

Earlier quoted context omitted.

> why are we concluding that bigger models and more data = more hallucination? That’s not what your quotes said. They said bigger models = plateau in intelligence, nothing about more data or increased hallucinations The relevant quote for what you’re talking about would be: > It’s been proven that when a model is trained on large volumes of highly factual and non-theoretical data, it learns to always have an answer.…

I find these internet arguments talking about LLMs as if they are trained by reading the internet to be wild. Yes, pretraining still exists. But for the past few years, pretraining by reading the internet is just the initial bootstrapping of LLM training. The RL training they get from bespoke training data, with very very different characteristics than what these armchair analyses claim, dominates these days.

In the last few weeks Claude (Sonnet) has told me “I don’t know” 3 different times. That seems like the solution to hallucinations and it’s already happening.

Re: GPT-5.5 hallucinates 3x more than MIT-licensed GLM-5.2

#268
post #242

> it is clear that actual intelligence has plateaued significantly N=1, but I disagree strongly. I'm writing a hard-science science fiction story, and the physics of it is at (and frankly, beyond) my skillset. The story's plot has had to change over a dozen times as I realized errors in my application of physics in the story. Throughout, I've been reviewing the physics with LLMs, mainly Gemini 3.1 Pro Preview, but al…

Hah, I noticed the same thing writing fiction with fable. Most models seem to go into a sort of "storytelling mode" where they forget their PhD level smarts. I had a character who is doing repair on a satellite. Most models would give you a half-baked explanation with some technical terms - half of them right half of them wrong. Fable gave a description so deep that even I couldn't figure out what was going on and ha…

Nice to hear N=2. I'm really hoping Fable comes back soon.

In my case two people are making very-near-light-speed trips to a star 20-ish light years away. Originally, I had one leaving a month earlier and making the journey with a Lorentz factor of 40, while the protagonist takes the same trip at > 200.

The former experiences a trip of 6 months, the latter something like 25 days. And I wrote it as if that meant that the protagonist would get there months ahead. But both of them will take hours to a day over the time light takes, and the one who leaves a month earlier will get almost a month before.

That error sat in my manuscript for two months of back and forth with other models. Fable found it on the first go.

LMK if you want to trade manuscripts!

Re: GPT-5.5 hallucinates 3x more than MIT-licensed GLM-5.2

#269

Earlier quoted context omitted.

Well the difference here is that you're overly simplifying complex biology and many other factors whereas llms are in fact actually simple mathematical models. As always, the devil lies in the details. Dismissing intricacies is a useful tool for daydreamers, not so much for engineers.

LLMs actually aren't simple Markov chains tho, your also simplifying. and LLMs trained with RLVR aren't just optimized over the space of functions (like gpt2 was), they're optimized over the space of programs (programs under some length). You find the ideal algorithm that can do the task you need it to.

RLVR is a process which updates the Markov chain

Re: GPT-5.5 hallucinates 3x more than MIT-licensed GLM-5.2

#270
This reminds me of the Missing Dollar Riddle [1], where the listener is deliberately put on a wrong thinking path, to fool it.

With your own logical thinking you might never come to this confusion, and if you never heard this riddle before, you might be tricked by it.

But as we grow in life, and get experience, we learn about these riddles and aren't fooled as easily anymore.

Maybe it'll work like that for LLMs too?

[1]: https://en.wikipedia.org/wiki/Missing_dollar_riddle

Post reply on HN