Live data from Hacker News

GPT-5.5 hallucinates 3x more than MIT-licensed GLM-5.2

arrowtsx.dev

251–260 of 318 posts

Re: GPT-5.5 hallucinates 3x more than MIT-licensed GLM-5.2

#251

Earlier quoted context omitted.

Which is why rubrics as rewards are used.

still cost prohibitive.

Yes, which is why for some things I've gotten paid as much as $1500 per training example generated.

AI labs don't care about cost prohibitive.

Re: GPT-5.5 hallucinates 3x more than MIT-licensed GLM-5.2

#252

Earlier quoted context omitted.

> Noone is really caring about hallucinations on point facts these days though, it is much more about complex reasoning tasks. The boundary is pretty thin there though. E.g., Gemini recently told me that a certain papers claims that two frameworks are mathematically equivalent, while the paper shows the opposite, and yesterday Google's AI overview told me that no World Cup matches were scheduled for that day despite…

That is a great example of the kind of thing they're paying people to create as training data. You write the prompt, and then write rubrics to judge the responses, and you found something the model failed at. Congratulations, you just earned $500, now do it again.

Not the worst way to make money, but if internet-scale data were not enough to reduce errors to a somewhat tolerable margin, how much data do they hope to collect in this manner?

Re: GPT-5.5 hallucinates 3x more than MIT-licensed GLM-5.2

#254

Earlier quoted context omitted.

That is a great example of the kind of thing they're paying people to create as training data. You write the prompt, and then write rubrics to judge the responses, and you found something the model failed at. Congratulations, you just earned $500, now do it again.

Not the worst way to make money, but if internet-scale data were not enough to reduce errors to a somewhat tolerable margin, how much data do they hope to collect in this manner?

Right now, this is a 10-figure run rate industry.

They are generating a lot of this. Also remember it's not just quantity, it's roughly active learning - they're paying for training data that's at the classification boundary, which is way more valuable.

I have gotten offers for contracts for full time jobs at high rates with AI labs to do this.

Meta has reallocated a lot of their full time SWE staff to do this.

All of this has rapidly accelerated within the last 6 months, who knows far it will go, if someone showed me a Kalshi bet that 10% of the college educated population of the US would be doing this as their primary job by the end of 2027, I wouldn't have the guts to bet against it.

10% of physicians' earnings doing this? Yeah that would totally track.

It doesn't seem like there's a limit. There's a shortage of GPUs and TSMC can only scale up so fast, so the AI labs found something else to spend money on.

Re: GPT-5.5 hallucinates 3x more than MIT-licensed GLM-5.2

#255

> it is clear that actual intelligence has plateaued significantly. > Moving forward, the industry cannot continue to train bigger and bigger models since their intelligence not only plateaus but often will get worse These are wild claims - why are we concluding that bigger models and more data = more hallucination? That’s actually the opposite of what’s been happening over the last couple years. Some models may stil…

> why are we concluding that bigger models and more data = more hallucination? That’s not what your quotes said. They said bigger models = plateau in intelligence, nothing about more data or increased hallucinations The relevant quote for what you’re talking about would be: > It’s been proven that when a model is trained on large volumes of highly factual and non-theoretical data, it learns to always have an answer.…

#2 is not that surprising from first principles if the way you made the bigger model was by feeding it poorer quality training data because it’s the only way you can get enough

Re: GPT-5.5 hallucinates 3x more than MIT-licensed GLM-5.2

#256

Earlier quoted context omitted.

That is informative, I was suspecting that is how models improve their performance on some convoluted "non-googlabe" benchmarks like SimpleBench, that is how, they just got the taste of those those questions from publicly available samples and then hired people to generate similar questions and provide answers for them. I wonder if extracting those static reasoning chains make sense given a Rich Sutton's "The Bitter…

There is one level that these training data give examples of specific static reasoning chains. Given exposure to enough reasoning chains, with training data that is designed around adversarial reasoning and teaching models to reason, these types of training data might be key to teaching models to reason beyond what they could gather from static data.

> these types of training data might be key to teaching models to reason beyond what they could gather from static data.

I was under impression that every time LLMs try to be truly novel and they need to assume things in the area where they didn't have enough data points that there were trained on, results are not good, has that changed?

Re: GPT-5.5 hallucinates 3x more than MIT-licensed GLM-5.2

#257

Earlier quoted context omitted.

There is one level that these training data give examples of specific static reasoning chains. Given exposure to enough reasoning chains, with training data that is designed around adversarial reasoning and teaching models to reason, these types of training data might be key to teaching models to reason beyond what they could gather from static data.

> these types of training data might be key to teaching models to reason beyond what they could gather from static data. I was under impression that every time LLMs try to be truly novel and they need to assume things in the area where they didn't have enough data points that there were trained on, results are not good, has that changed?

If LLMs were already good at it, the AI labs wouldn't be paying this insane amount of money for people to generate training data to teach them.

Re: GPT-5.5 hallucinates 3x more than MIT-licensed GLM-5.2

#258

Earlier quoted context omitted.

Yeah not only is it totally unsubstantiated, the benchmarks are getting less useful to really show the difference between these models. Big model smell is still a thing and GLM 5.2 while impressive is not Fable class. Here is something I would like people to chew on. Perhaps the smartest researchers in the world across multiple labs know more about this than we do? Perhaps they are aware of issues like the data wall…

Are the smartest researchers in the world out there saying there isn't a wall? I don't know of any people doing the actual R&D who frequently make outrageous claims.

The entirety of Anthropic believe ai is going to eat everything, not just software, and result in major societal disruption within a year. They do not have a sliver of a doubt on this. Article has no idea, is completely wrong.

Re: GPT-5.5 hallucinates 3x more than MIT-licensed GLM-5.2

#259

Earlier quoted context omitted.

Not the worst way to make money, but if internet-scale data were not enough to reduce errors to a somewhat tolerable margin, how much data do they hope to collect in this manner?

Right now, this is a 10-figure run rate industry. They are generating a lot of this. Also remember it's not just quantity, it's roughly active learning - they're paying for training data that's at the classification boundary, which is way more valuable. I have gotten offers for contracts for full time jobs at high rates with AI labs to do this. Meta has reallocated a lot of their full time SWE staff to do this. All o…

Yes, they do have money to burn, and this will bring some improvements for sure, but active learning has never really worked out, has it? And even 10% of the educated population doing this for, like, 50 years is not that much data, while normally each accuracy percentage is more and more data-expensive.
Post reply on HN