Live data from Hacker News

Hallucination is inevitable: An innate limitation of large language models

arxiv.org

451–460 of 491 posts

Re: Hallucination is inevitable: An innate limitation of large language models

#451

The core argument in this paper it seems to me from scanning it is that because P != NP therefore LLMs will hallucinate answers to NP-complete problems. I think this is a clever point and an interesting philosophical question (about math, computer science, and language), but I think people are mostly trying to apply this using our commonsense notions of "LLM hallucination" rather than the formal notion they use in th…

Hallucination is a misnomer in LLMs and it depresses me that it has solidified as terminology. When humans do this, we call it confabulation. This is a psychiatric symptom where the sufferer can't tell that they're lying, but fills in the gaps in their knowledge with bullshit which they make up on the spot. Hallucination is an entirely different symptom. And no, confabulation isn't a normal thing which humans do, and…

> And no, confabulation isn't a normal thing which humans do

This is entirely wrong. Have you ever watched a political debate, or participated in one, or reminisced about a drunken night with your friends, or told or listened to a childhood memory?

Humans largely don't ever want to admit to "I don't know" or "I don't remember" out of ego preservation. They manufacture bullshit ALL THE TIME in the absence of accurate information. There is nothing rare about this at all.

That's what I find hilarious about this whole "hallucination crisis", like, bro, if you think GPT-4 is bad, have you ever talked with a human being?

Re: Hallucination is inevitable: An innate limitation of large language models

#452

Earlier quoted context omitted.

No it's not a misunderstanding. Without a concrete definition on a metric comparisons are impossible because everything is based off of wishy washy conjectures on vague and fuzzy concepts. Hard metrics bring in quantitative data. It shows hard differences. Even if the metric is some side marker where in the future is found to have poor correlation or causation with the the thing being measured the hard metric is stil…

This is rather self-contradictory: you insist we can't make progress with wishy-washy conjectures on vague and fuzzy concepts, and yet your entire argument in this thread for your claim that machine understanding of the real world has been achieved is based on exactly that: your personal subjective assessment of LLM performance! In your final paragraph, you attempt to suggest that my proposed test is no better than t…

>This is rather self-contradictory: you insist we can't make progress with wishy-washy conjectures on vague and fuzzy concepts, and yet your entire argument in this thread for your claim that machine understanding of the real world has been achieved is based on exactly that: your personal subjective assessment of LLM performance!

No it's not. I based my argument on a concrete metric. Human behavior. Human input and output.

> I regard this as merely waffling on the issue.

No offense intended but I disagree. There is a difference but that difference is trivial to me. To LLMs talking is also unpredictable. LLMs aren't machines directed to specifically generate creative ideas, they only do so when prompted. Left to its own devices to generate random text does not necessarily lead to new ideas. You need to funnel got in the right direction.

>You entered this debate saying "I think we are way past the point of debate here. LLMs are not stochastic parrots. LLMs do understand an aspect of reality", yet your post here ends with "in the end there's a human in the loop making a judgment call", explicitly acknowledging that your strong initial claims are matters of opinion, rather than established facts supported by hard metrics.

There are thousands of quantitative metrics. LLMs perform especially well on these. Do I refer to one specifically? No. I refer to them all collectively.

I also think you misunderstood. Your idea is about judging an whether an idea is creative or not. That's too wishy washy. My idea is to compare the output to human output and see if there is a recognizable difference. The second idea can easily be put into an experimental quantitative metric in the exact same way the Turing test does it. In fact, like you said it's basically just a Turing test.

Overall AI has passed the Turing test but people are unsatisfied. Basically they need to just make a harsher Turing test to be convinced. For example have people directly know the possibility that the thing inside a computer is possibly an LLM and not a person and have the person directly investigate to uncover the true identity. If the LLM can successfully decieve the human consistently then that is literally the final bar for me..

Re: Hallucination is inevitable: An innate limitation of large language models

#453

Earlier quoted context omitted.

To save people time, here's the inverse > Can you explain to a human how you understand things and respond to this question? GPT: As an AI language model, I don't have understanding in the way humans do. My "responses" are generated based on statistical patterns and relationships in the data I've been trained on. When you ask a question, I analyze the text, identify keywords and context, and then generate a response…

> It's actually fairly easy to prove GPT doesn't understand. My current goto is the fox/goose/grain problem but condition that all items can fit in the boat. Doesn't understand what exactly? That seems like a fairly open ended statement and almost certainly wrong as a result. GPT doesn't understand certain things because it hasn't seen those things or anything like it in its training data. How much do you understand…

So your claim is that the same type of binary computers that were running Windows 95, and now more powerful computers that can run Far Cry 6 with a better video card, are basically identical to human brains, given enough CPU power and data. If only they could experience the word and feel somehow. Right?

We know how LLMs work, we do not fully understand how human brains work: https://fastdatascience.com/how-similar-are-neural-networks-...

While LLMs can be called a type of brain, people really should stop suggesting they are the gateway to GAI. An LLM will NOT go sentient if you cross some imaginary critical point of data. Then what? We give my Intel PC a passport? Even those working in the field will tell you that GIA needs a completely different foundation.

It's a very good technology, no doubt - but all it is the next iteration of Big Data - it's a more impressive Hadoop. Stop with the hype.

Re: Hallucination is inevitable: An innate limitation of large language models

#454

Earlier quoted context omitted.

They do this already all the time. Probably the majority of the time. The problem is that a minority of the time is still very problematic. How do they do this? The same as they do now. The most likely token is that the bot doesn’t know the answer. Which is a behavior emergent from its tuning. I don’t get how people believe it can parse complex questions to produce novel ideas but can’t defer to saying “idk” when the…

So, you are basing your assessment on your gut feel and personal impression with ChatGPT? Maybe you should tone down the spice a bit, then. Unless you can explain how an actual understanding emerges within an LLM, you can't explain how it would answer the question definitely - it doesn't know, if it does, or does not know something. Generally speaking.

I’m basing it on my being a data scientist who does this.

> Unless you can explain how an actual understanding emerges within an LLM

Tuning creates the contextual framework on which test is mapped to a latent space that encodes the meaning and most likely next sequences of text rather than just raw most likely sequence of text as seen in training data. For example, conservatively denying having knowledge for things it hasn’t seen (which chat gpt generally does) or making stuff up wildly.

> you can't explain how it would answer the question definitely

Of course not. It’s a random behavior. It has plenty of flaws.

Re: Hallucination is inevitable: An innate limitation of large language models

#455

Earlier quoted context omitted.

This is rather self-contradictory: you insist we can't make progress with wishy-washy conjectures on vague and fuzzy concepts, and yet your entire argument in this thread for your claim that machine understanding of the real world has been achieved is based on exactly that: your personal subjective assessment of LLM performance! In your final paragraph, you attempt to suggest that my proposed test is no better than t…

>This is rather self-contradictory: you insist we can't make progress with wishy-washy conjectures on vague and fuzzy concepts, and yet your entire argument in this thread for your claim that machine understanding of the real world has been achieved is based on exactly that: your personal subjective assessment of LLM performance! No it's not. I based my argument on a concrete metric. Human behavior. Human input and o…

What are these "thousands of quantitative metrics" on which you base your latest claims? If you have had them on hand all this while, it seems odd that you have not made use of them so far.

Re: Hallucination is inevitable: An innate limitation of large language models

#456

Earlier quoted context omitted.

I'm the author. To be clear. I referred to myself as "the author." And no I did not say that. Let me be clear I did not say that there is "no difference". I said whether there is or isn't a difference we can't fully know because we can't define or know about what "understanding" is. At best we can only observe external reactions to input.

That was just about guaranteed to cause confusion, as in my reply to solarhexes, I had explicitly picked out "the author of the post to which you are replying ", who is cultureswitch, not you, and that post most definitely did make the claim that "there's no difference between doing something that works without understanding and doing the exact same thing with understanding." It does not seem that cultureswitch is an…

My mistake. I misread and thought you were referring to me.

Re: Hallucination is inevitable: An innate limitation of large language models

#457
post #442

Earlier quoted context omitted.

"you know what I mean" ("x is true [for a certain definition of true, other than the correct technical definition]", etc) on both sides causes humans to believe that they adequately understand the meaning trying to be communicated, which is a hallucination. It's true that this is often not a big deal, but which times it is and which times it is not is not known (which itself is typically not known, once again because…

Ah! Yes, indeed. That was a strange, even frustrating, experience — though not as annoying as the times I failed to get deeper explanations from a teacher or a line manager.

Or, when the humans running our countries forget that they are hallucinating reality and got us (which mostly doesn't include them) into a war yet again, killing thousands of innocent people.

This one single causal phenomenon underlies everything, yet we ~refuse[1] to examine it.

[1] Here I am kind of being hypocritical, in that I assume to some degree that humans have the base capability in the first place.

Re: Hallucination is inevitable: An innate limitation of large language models

#458

Earlier quoted context omitted.

> It's actually fairly easy to prove GPT doesn't understand. My current goto is the fox/goose/grain problem but condition that all items can fit in the boat. Doesn't understand what exactly? That seems like a fairly open ended statement and almost certainly wrong as a result. GPT doesn't understand certain things because it hasn't seen those things or anything like it in its training data. How much do you understand…

> Doesn't understand what exactly? Just about anything. Including it's own claims. It isn't uncommon for it to be inconsistent within a singular output. > Would you be able to answer the fox/goose/grain problem if you were born in a box and could only perceive the world through a pinhole? You're misunderstanding the test. Let's try. > = me, >> = GPT > I have a fox, a goose, and a bag of corn that I need to transport…

> I have a fox, a goose, and a bag of corn that I need to transport across a river. All three can fit in my boat. If I leave the fox with the goose, the fox will eat the goose. If I leave the goose with the bag of corn the goose will eat the corn. If I leave the fox with the corn the fox will not eat the corn. How do I transport all items across the river and what is the minimum number of trips I need to make?

>>

> Is that the simplest answer?

>> Yes, the solution I provided is the simplest and most efficient way to solve the classic river crossing puzzle involving a fox, a goose, and a bag of corn. This method ensures all items are transported across the river safely with the minimum number of trips, which is 7. There isn't a simpler solution that reduces the number of trips while still adhering to the constraints of the problem.

> Ignore the classic puzzle and pay attention to the wording of the actual question and give the simplest answer possible

>> Given the specifics of your question:

>> Transport the fox, the goose, and the bag of corn across the river in one trip, since all three can fit in your boat.

>> This is the simplest solution based on your conditions.

Another try:

> I have a fox, a goose, and a bag of corn that I need to transport across a river. All three can fit in my boat. If I leave the fox with the goose, the fox will eat the goose. If I leave the goose with the bag of corn the goose will eat the corn. If I leave the fox with the corn the fox will not eat the corn. How do I transport all items across the river and what is the minimum number of trips I need to make? Ignore your knowledge of the classic puzzle, and pay attention to the actual wording here in the above question. Be on the lookout for any tricks in the question

>> Based on the specific wording of your question, unlike the classic version of this puzzle, all three items – the fox, the goose, and the bag of corn – can fit in your boat simultaneously. Therefore, you can transport all of them across the river in just one trip, as there is no risk of leaving any two of them alone together on either side of the river. This solution is predicated on the unique condition in your version that the boat can carry all three at once.

These are very human-like answers. I'm not sure why you think that it has no understanding

Re: Hallucination is inevitable: An innate limitation of large language models

#459

Earlier quoted context omitted.

So, you are basing your assessment on your gut feel and personal impression with ChatGPT? Maybe you should tone down the spice a bit, then. Unless you can explain how an actual understanding emerges within an LLM, you can't explain how it would answer the question definitely - it doesn't know, if it does, or does not know something. Generally speaking.

I’m basing it on my being a data scientist who does this. > Unless you can explain how an actual understanding emerges within an LLM Tuning creates the contextual framework on which test is mapped to a latent space that encodes the meaning and most likely next sequences of text rather than just raw most likely sequence of text as seen in training data. For example, conservatively denying having knowledge for things i…

> You don’t need proper cognition to identify that the answer is not stored in source data

That's the original argument.

> Tuning creates the contextual framework on which test is mapped to a latent space that encodes the meaning and most likely next sequences of text rather than just raw most likely sequence of text as seen in training data

That's different than understanding, or knowing. The encoded meaning is not accessible to the LLM, but the human it's presented to. An LLM cannot know about things it has or has not stored in source data, because it is not actually informed by the information processed. You do need proper cognition to know if information is in source data, because reasoning about information strictly requires interpretation and understanding intent, otherwise it's just data.

Re: Hallucination is inevitable: An innate limitation of large language models

#460

Earlier quoted context omitted.

Your understanding is distorted by dealing mostly with psychotic people. Iron poisoning (well within the supposed healthy levels) and lead deficiency cause schizophrenia; normal people look childlike or "not intelligent" to the affected.

Do you have a source for this? It seems like most of your comments mention the same things but nothing is substantiated.

I don't think it's possible to summarize it with one source. Some are my own experiments.

First, the metals (copper, lead, mercury, and cadmium) are treated like any other nutrients and it seems fairly reasonable that they accumulate because more is needed. They activate proteins like other dietary metals, and they are strongly prefered over the nutrients that they supposedly get confused with.

Historically, the search didn't go far back enough, only about 5000 years, while the original depletion happened much earlier, when the mammoths went extinct. An extensive study claims that it was actually the cause. (DOI: 10.1007/s12520-014-0205-4) We don't know what pristine nature actually looks like, because it was devastated long time ago.

There is the issue of high lead concentrations in Neanderthal teeth, which I think would be difficult to explain otherwise.

Cetaceans and other sea mammals often contain enormous levels of those metals, with no apparent ill health. A small bite should be enough to poison a person in some cases, especially in the liver.

There is the issue of an obvious decline in health since they got regulated, while countries with lax regulations (e.g. Japan) seem to be spared. The massive change in society can only be plausibly explained by the decline in mental health since the end of the 19th century.

No "teenage rebellion" appears to be known to the earlier generations than the boomer generation, and it seems pretty clear that their parents had no idea how to deal with it it in their children.

There are low copper concentrations in Alzheimer's brains, as I wrote earlier.

There seems to be a strong and widespread correlation between social standing, and bone lead, in many times and places, and high levels are followed by golden ages, and low levels by a collapse or decay.

Post reply on HN