Live data from Hacker News

LLMs Will Always Hallucinate, and We Need to Live with This

arxiv.org

241–250 of 274 posts

Re: LLMs Will Always Hallucinate, and We Need to Live with This

#241

Earlier quoted context omitted.

If you ask about opinions, sure. Because there are no "true" opinions. If you ask about the capital of France, any answer but "Paris" is objectively wrong, whether given by a human or LLM.

Paris has not always been the capital of France. Many other cities around France have been capital. https://en.wikipedia.org/wiki/List_of_capitals_of_France There's practically no subject you could bring up that an LLM wouldn't "hallucinate" or give "wrong" information about given that garbage in -> garbage out , and LLMs are trained on all the garbage (as well as too many facts) they've been able to scrape. The LLM…

But if you ask "what is the capital of France" (not what it has been, but what it is), there is actually only one correct answer. "Capital" has a definition, and the headquarters of the government of France has a definite location. Sure, some French citizens will give a different answer. Some people will say the earth is flat, too. They are wrong.

Re: LLMs Will Always Hallucinate, and We Need to Live with This

#242

Earlier quoted context omitted.

If you ask about opinions, sure. Because there are no "true" opinions. If you ask about the capital of France, any answer but "Paris" is objectively wrong, whether given by a human or LLM.

Paris has not always been the capital of France. Many other cities around France have been capital. https://en.wikipedia.org/wiki/List_of_capitals_of_France There's practically no subject you could bring up that an LLM wouldn't "hallucinate" or give "wrong" information about given that garbage in -> garbage out , and LLMs are trained on all the garbage (as well as too many facts) they've been able to scrape. The LLM…

This got real tedious...

Re: LLMs Will Always Hallucinate, and We Need to Live with This

#243

Earlier quoted context omitted.

You can simulate a NAND gate using balls rolling down a specially designed wood board. In theory you could construct a giant wood board with billions and billions of balls that would implement the inference step of an LLM. Do you see these balls rolling down a wood board as a form of interiority/subjective experience? If not, then why do you give it to electric currents in silicon? Just because it's faster?

Your point of disagreement is the _medium_ of computation? The same point can be made about neurons. Do you think you could have the same kind of cognitive processes you have now if you were thinking 1000x slower than you do? Speed of processing matters, especially when you have time bounds on reaction, such in real life. Another problem with balls would be the necessity of perception, that you can't really do with b…

> Your point of disagreement is the _medium_ of computation?

No. My point is that we should not impute interiority onto computation. The medium is a thought experiment meant to stimulate your intuition that computational sophistication does not imply interiority. Unless you do think that the current state of balls rolling down a board does entail a conscious experience?

To adjust the speed, just imagine the balls rolling down in 100000000x time. I don't know where you stand, but I still don't think there is a conscious experience on the board.

Re: LLMs Will Always Hallucinate, and We Need to Live with This

#244
post #146

Earlier quoted context omitted.

LLMs do now have a concept of truth now since much of the RLHF is focused on making them more accurate and true. I think the problem is that humanity has a poor concept of truth. We think of most things as true or not true when much of our reality is uncertain due to fundamental limitations or because we often just don't know yet. During covid for example humanity collectively hallucinated the importance of disinfect…

> humanity collectively hallucinated the importance of disinfecting groceries for awhile I reject this history. I homeschooled my kids during covid due to uncertainty and even I didn't reach that level, and nor did anyone I knew in person. A very tiny number who were egged on by some YouTubers did this, including one person I knew remotely. Unsurprisingly that person was based in SV.

I did not do this personally but I know a number of people (blue state liberal city folk) I don’t think it was that unusual.

Re: LLMs Will Always Hallucinate, and We Need to Live with This

#245

Earlier quoted context omitted.

Paris has not always been the capital of France. Many other cities around France have been capital. https://en.wikipedia.org/wiki/List_of_capitals_of_France There's practically no subject you could bring up that an LLM wouldn't "hallucinate" or give "wrong" information about given that garbage in -> garbage out , and LLMs are trained on all the garbage (as well as too many facts) they've been able to scrape. The LLM…

This got real tedious...

lame response.

Re: LLMs Will Always Hallucinate, and We Need to Live with This

#246

Earlier quoted context omitted.

Exactly this, I've been saying this since the beginning. Every response is a hallucination - a probabilistic string of words divorced from any concept of truth or reality. By total coincidence, some hallucinations happen to reflect the truth, but only because the training data happened to generally be truthful sentences. Therefore, creating something that imitates a truthful sentence will often happen to also be trut…

I think you're going too far here. > By total coincidence, some hallucinations happen to reflect the truth, but only because the training data happened to generally be truthful sentences. It's not a "total coincidence". It's the default. Thus, the model's responses aren't "divorced from any concept of truth or reality" - the whole distribution from which those responses are pulled is strongly aligned with reality. (W…

> Even the most blatant lies, even all of fiction writing, they're all incorrect or fabricated only at the surface level - the whole thing, accounting for the utterance, what it is about, the meanings, the words, the grammar - is strongly correlated with truth and reality.

I would reject this pretty firmly. As you said, people write whole novels about imagined worlds and people about magic or technology or whatever that doesn't or can't exist. The LLM may understand what words mean and "know" how to string them together into a meaningful and grammatical sentence, but that's entirely different than a truthful sentence.

Truth requires some mechanism of fact finding, or chains of evidence, or admitting when those chains don't exist. LLMs have nothing like that.

Re: LLMs Will Always Hallucinate, and We Need to Live with This

#247

Earlier quoted context omitted.

Paris has not always been the capital of France. Many other cities around France have been capital. https://en.wikipedia.org/wiki/List_of_capitals_of_France There's practically no subject you could bring up that an LLM wouldn't "hallucinate" or give "wrong" information about given that garbage in -> garbage out , and LLMs are trained on all the garbage (as well as too many facts) they've been able to scrape. The LLM…

But if you ask "what is the capital of France" (not what it has been, but what it is ), there is actually only one correct answer. "Capital" has a definition, and the headquarters of the government of France has a definite location. Sure, some French citizens will give a different answer. Some people will say the earth is flat, too. They are wrong .

But we're talking about LLMs in this thread, and the example I used of French citizens not always saying Paris is the capital of France is just an example of how topics can be subjective. If you have something pertinent to the LLM discussion, then please reply.

Re: LLMs Will Always Hallucinate, and We Need to Live with This

#248

Earlier quoted context omitted.

Exactly this, I've been saying this since the beginning. Every response is a hallucination - a probabilistic string of words divorced from any concept of truth or reality. By total coincidence, some hallucinations happen to reflect the truth, but only because the training data happened to generally be truthful sentences. Therefore, creating something that imitates a truthful sentence will often happen to also be trut…

In other words, all models are wrong, but some are useful.

Precisely, I'm glad someone picked up on the reference!

Re: LLMs Will Always Hallucinate, and We Need to Live with This

#249

Earlier quoted context omitted.

Exactly this, I've been saying this since the beginning. Every response is a hallucination - a probabilistic string of words divorced from any concept of truth or reality. By total coincidence, some hallucinations happen to reflect the truth, but only because the training data happened to generally be truthful sentences. Therefore, creating something that imitates a truthful sentence will often happen to also be trut…

Ok, but I think it would be more productive to educate people that LLMs have no concept of truth rather than insist they use the term "hallucinate" in an unintuitive way.

If people already understand what "hallucination" means, then I think it's perfectly intuitive and educational to say that, actually, the LLM is always doing that, just that some of those hallucinations happen to coincidentally describe something real.

We need to dispell the notion that the LLM "knows" the truth, or is "smart". It's just a fancy stochastic parrot. Whether it's responses reflect a truthful reality or a fantasy it made up is just luck, weighted by (but not constrained to) its training data. Emphasizing that everything is a hallucination does that. I purposefully want to reframe how the word is used and how we think about LLMs.

Re: LLMs Will Always Hallucinate, and We Need to Live with This

#250

> By establishing the mathematical certainty of hallucinations, we challenge the prevailing notion that they can be fully mitigated Having a mathematical proof is nice, but honestly this whole misunderstanding could have been avoided if we'd just picked a different name for the concept of "producing false information in the course of generating probabilistic text". "Hallucination" makes it sound like something is goi…

It’s also problematic to suggest that the problem is unsolvable. It’s unsolvable inside the LLM. It’s 100% solvable within the product.
Post reply on HN