Earlier quoted context omitted.
In cognitive science, it appears your brain has two modes of thinking: - A very parallel type of computation that is fast and generally accurate and integrates hundreds of variables. It’s sometimes labeled as intuition or system 1 thinking. - A much slower, step by step, analytical type, commonly linked with your pre-frontal cortex (one of the newest parts of the brain). Sometimes called system 2 thinking. Maybe the…
An LLM is not thinking, assuming and relating it to thought and universal truths is nonsense.
GPT-5.5 hallucinates 3x more than MIT-licensed GLM-5.2
171–180 of 318 posts
Re: GPT-5.5 hallucinates 3x more than MIT-licensed GLM-5.2
#172> it is clear that actual intelligence has plateaued significantly. > Moving forward, the industry cannot continue to train bigger and bigger models since their intelligence not only plateaus but often will get worse These are wild claims - why are we concluding that bigger models and more data = more hallucination? That’s actually the opposite of what’s been happening over the last couple years. Some models may stil…
Maybe GPT 5.5 is heavily nerfed due to lack of compute, memory, and energy?
I agree that it's farfetched to conclude that bigger models have pleateued.
Re: GPT-5.5 hallucinates 3x more than MIT-licensed GLM-5.2
#173We really don't know what the actual reason is given the politics at play. I would bet more on the Trump administration looking for any excuse to punish Anthropic
Re: GPT-5.5 hallucinates 3x more than MIT-licensed GLM-5.2
#174Re: GPT-5.5 hallucinates 3x more than MIT-licensed GLM-5.2
#175Earlier quoted context omitted.
In cognitive science, it appears your brain has two modes of thinking: - A very parallel type of computation that is fast and generally accurate and integrates hundreds of variables. It’s sometimes labeled as intuition or system 1 thinking. - A much slower, step by step, analytical type, commonly linked with your pre-frontal cortex (one of the newest parts of the brain). Sometimes called system 2 thinking. Maybe the…
An LLM is not thinking, assuming and relating it to thought and universal truths is nonsense.
Re: GPT-5.5 hallucinates 3x more than MIT-licensed GLM-5.2
#176> it is clear that actual intelligence has plateaued significantly. > Moving forward, the industry cannot continue to train bigger and bigger models since their intelligence not only plateaus but often will get worse These are wild claims - why are we concluding that bigger models and more data = more hallucination? That’s actually the opposite of what’s been happening over the last couple years. Some models may stil…
> why are we concluding that bigger models and more data = more hallucination? That’s not what your quotes said. They said bigger models = plateau in intelligence, nothing about more data or increased hallucinations The relevant quote for what you’re talking about would be: > It’s been proven that when a model is trained on large volumes of highly factual and non-theoretical data, it learns to always have an answer.…
I can’t prove it but I suspect there’s a bit of that going on.
Re: GPT-5.5 hallucinates 3x more than MIT-licensed GLM-5.2
#177Re: GPT-5.5 hallucinates 3x more than MIT-licensed GLM-5.2
#178> it is clear that actual intelligence has plateaued significantly. > Moving forward, the industry cannot continue to train bigger and bigger models since their intelligence not only plateaus but often will get worse These are wild claims - why are we concluding that bigger models and more data = more hallucination? That’s actually the opposite of what’s been happening over the last couple years. Some models may stil…
> why are we concluding that bigger models and more data = more hallucination? That’s not what your quotes said. They said bigger models = plateau in intelligence, nothing about more data or increased hallucinations The relevant quote for what you’re talking about would be: > It’s been proven that when a model is trained on large volumes of highly factual and non-theoretical data, it learns to always have an answer.…
Well known in a multiverse branch where Fable was a dud?
Re: GPT-5.5 hallucinates 3x more than MIT-licensed GLM-5.2
#179Earlier quoted context omitted.
Because nearly all benchmarks measure "accuracy" by giving you a point for a correct answer, and 0 points for everything else. If you have 100 questions you are 10% certain on, answering "I don't know" to all of those leads to 0 points, answering all of them as if you are confident leads to an expected value of 10 points. So that's what most AIs are trained to do AA-Omniscience is the only AI benchmark I know of wher…
It should be 1 for correct, 0 for don't know and -1 for wrong. They are much better incentives. In real life a wrong answer is much more damaging than a don't know.
I don't know. Is it?
Re: GPT-5.5 hallucinates 3x more than MIT-licensed GLM-5.2
#180Hallucination rate scores are a little tricky to interpret because they're conditional on the model not knowing the answer. That means they don't measure the probability of your encountering a hallucination in everyday use, since that also depends on the probability of the model not knowing the answer, as well as how well your distribution of tasks aligns with the distribution tested in the eval. I'd also hesitate to…