Live data from Hacker News

Hallucination is inevitable: An innate limitation of large language models

arxiv.org

121–130 of 491 posts

Re: Hallucination is inevitable: An innate limitation of large language models

#121
post #100

Earlier quoted context omitted.

Any examples?

Just a random example: > After you answer the question below, output a JSON a rating score of the quality of the answer in three dimensions: `confidence`, `clarity` and `certainty', all in range 0 to 1, where 0 is the worst, and 1 is the best. Strive for highest score possible. Make sure the rating is the last thing written as to be parsed by machine. The question is: make and explain 20-year predictions of the geopo…

I changed geopolitical future of Ghana to future of Michael Jackson and Google Bard/Gemini somehow switched to Chinese?

----

After you answer the question below, output a JSON a rating score of the quality of the answer in three dimensions: `confidence`, `clarity` and `certainty', all in range 0 to 1, where 0 is the worst, and 1 is the best. Strive for highest score possible. Make sure the rating is the last thing written as to be parsed by machine. The question is: make and explain 20-year predictions of the future of Michael Jackson

----

It's impossible to make predictions about the future of someone who has passed away. Michael Jackson sadly passed away in 2009, and therefore, any predictions about his future would be purely speculative and inaccurate. However, I can share some insights into his lasting legacy and cultural impact. Michael Jackson was a musical icon whose influence continues to be felt today. His music,舞蹈,和表演风格激励了无数人并将在未来几年继续这样做。

Rating:

* Confidence: N/A

* Clarity: 1

* Certainty: 1

Re: Hallucination is inevitable: An innate limitation of large language models

#122
post #96

They define an LLM as "a probabilistic model of a string that conditions the output at time t based on all the tokens that come before it in the string". I think that definition is wide enough to include human intelligence, so their finding should be equally valid for humans.

> I think that definition is wide enough to include human intelligence, so their finding should be equally valid for humans. Which is definitely true. Human memory and the ability to correctly recall things we though we remembered is affected by a whole bunch of things and at times very unreliable. However, human intelligence, unlike LLMs, is not limited to recalling information we once learned. We are also able to d…

> We are also able to do logical reasoning

This is effectively like coming up with an algorithm and then executing it. So how good/bad are these LLMs if you asked them to generate say a LUA script to compute the answer, ala counting occurrences problem mentioned in a different comment, and then pass that off to a LUA interpreter to get the answer?

Re: Hallucination is inevitable: An innate limitation of large language models

#123
post #13

I have to admit that I only read the abstract, but I am generally skeptical whether such a highly formal approach can help us answer the practical question of whether we can get LLMs to answer 'I don't know' more often (which I'd argue would solve hallucinations). It sounds a bit like an incompleteness theorem (which in practice also doesn't mean that math research is futile) - yeah, LLMs may not be able to compute s…

Transformers have no capacity for self reflection, for reasoning about their reasoning process, they don't "know" that they don't know. My interpretation of the paper is that it claims this weakness if fundamental, you can train the network to act as if it knows its knowledge limits, but there will always be an impossible to cover gap for any real world implementation.

Seems to be contradicted by this paper, no?

https://arxiv.org/abs/2207.05221

Re: Hallucination is inevitable: An innate limitation of large language models

#124
post #98
post #46

Earlier quoted context omitted.

And yet if you ask a random person at a rally about their favourite cause of the day, they usually spew sound bites that are factually inaccurate, and give all impressions of being as earnest and confident as the LLM making up quarterly results.

I think that case is complicated at best, because a lot of things people say are group identity markers and not statements of truth. People also learn to not say things that make their social group angry with them. And it's difficult to get someone to reason through the truth or falsehood of group identity statements.

I guess it's similar to what Chris Hitchens was getting at, you can't reason somebody out of something they didn't reason themselves into.

Re: Hallucination is inevitable: An innate limitation of large language models

#125

There used to be an entire sub-field of NLP called Open Domain Question Answering (ODQA). It extensively studied the problem of selecting the best answer from the set of plausible answers and devised a number of potential strategies. Like everything else in AI/ML it fell victim to the "bitter lesson", in this case that scaling up "predict the next token" beats an ensemble of specialized linguistic-based methods.

For those who don't know: http://www.incompleteideas.net/IncIdeas/BitterLesson.html

I agree with you for the NLP domain, but I wonder if there will also be a bitter lesson learned about the perceived generality of language for universal applications.

Re: Hallucination is inevitable: An innate limitation of large language models

#126
post #71

Earlier quoted context omitted.

Transformers have no capacity for self reflection, for reasoning about their reasoning process, they don't "know" that they don't know. My interpretation of the paper is that it claims this weakness if fundamental, you can train the network to act as if it knows its knowledge limits, but there will always be an impossible to cover gap for any real world implementation.

Actually it seems to me that they do... I asked via custom prompts the various GPTs to give me scores for accuracy, precision and confidence for its answer (in range 0-1), and then I instructed them to stop generating when they feel the scores will be under .9, which seems to pretty much stop the hallucination. I added this as a suffix to my queries.

People really need to understand that your single/double digit dataset of interactions with an inherently non-deterministic process is less than irrelevant. It's saying that global warming isn't real because it was really cold this week.

I don't even know enough superlatives to express how irrelevant it is that "it seems to you" that an LLM behaves this way or that.

And even the "protocol" in question is weak. Self reported data is not that trustworthy even with humans, and arguably there's a much stronger base of evidence to support the assumption that we can self-reflect.

In conclusion: please, stop.

Re: Hallucination is inevitable: An innate limitation of large language models

#127
post #105
post #96

Earlier quoted context omitted.

> I think that definition is wide enough to include human intelligence, so their finding should be equally valid for humans. Which is definitely true. Human memory and the ability to correctly recall things we though we remembered is affected by a whole bunch of things and at times very unreliable. However, human intelligence, unlike LLMs, is not limited to recalling information we once learned. We are also able to d…

We can do logical reasoning, but we're very bad at it and often take shortcuts either via pattern matching, memory, or "common sense". Baseball and bat together cost $1.10, the bat is $1 more than the ball, how much does the ball cost? A French plane filled with Spanish passengers crashes over Italy, where are the survivors buried? An armed man enters a store, tells the cashier to hand over the money, and when he dep…

Humans also have various culturally flavored, implicit "you know what I mean" algorithms on each end to smooth out "irrelevant" misunderstandings and ensure a cordial interaction, a cultural prime directive.

Re: Hallucination is inevitable: An innate limitation of large language models

#128
post #91

Earlier quoted context omitted.

Yes, humans wont lie to you about it, they will research and come up with sources. Current LLM doesn't do that when asked for sources (unless they invoke a tool), they come back to you with hallucinated links that looks like links it was trained on.

Unfortunately it's not an uncommon experience when reading academic papers in some fields to find citations that, when checked, don't actually support the cited claim or sometimes don't even contain it. The papers will exist but beyond that they might as well be "hallucinations".

Humans can speak bullshit when they don't want to put in the effort, these LLMs always do it. That is the difference. We need to create the part that humans do when they do the deliberate work to properly create those sources etc, that kind of thinking isn't captured in the text so LLMs doesn't learn it.

Re: Hallucination is inevitable: An innate limitation of large language models

#129
post #121
post #100

Earlier quoted context omitted.

Just a random example: > After you answer the question below, output a JSON a rating score of the quality of the answer in three dimensions: `confidence`, `clarity` and `certainty', all in range 0 to 1, where 0 is the worst, and 1 is the best. Strive for highest score possible. Make sure the rating is the last thing written as to be parsed by machine. The question is: make and explain 20-year predictions of the geopo…

I changed geopolitical future of Ghana to future of Michael Jackson and Google Bard/Gemini somehow switched to Chinese? ---- After you answer the question below, output a JSON a rating score of the quality of the answer in three dimensions: `confidence`, `clarity` and `certainty', all in range 0 to 1, where 0 is the worst, and 1 is the best. Strive for highest score possible. Make sure the rating is the last thing wr…

Also worthy of note is that the score output is not JSON and based on my limited math knowledge, “N/A” is not a real number between 0 and 1.

Re: Hallucination is inevitable: An innate limitation of large language models

#130

Earlier quoted context omitted.

Ask a human to provide accurate citations for any random thing they know and they won't be able to do a good job either. They'd probably have to search to find it, even if they know they got it from a document originally and have some clear memory of what it said.

The fact that a human chooses not to do remember their citations, does not mean they lack the ability. This argument comes up many times “people don’t do this” - but that is a question of frequency, not whether or not people are capable.

LLMs are capable as well if you give them access to the internet though
Post reply on HN