Live data from Hacker News

Hallucination is inevitable: An innate limitation of large language models

arxiv.org

471–480 of 491 posts

Re: Hallucination is inevitable: An innate limitation of large language models

#471
post #147

Earlier quoted context omitted.

I think I agree with you (I even upvoted), but this might be an anthropomorphism. Back like 3-5 years ago, we already thought that about LLMs: They couldn't answer questions about what would fall when stuff are attached together in some non-obvious way, and the argument back then was that you had to /experience/ it to realize it. But LLMs have long fixed those kind of issues. The way LLMs "resolve" questions is very…

Think of it this way: Intelligent beings in the real world have a very complex built-in biological error function rooted in real world experiences: sensory inputs, feelings, physical and temporal limitations and so on. You feel pain, joy, fear, have a limited lifetime, etc. "AI" on the other hand only have an external error function, usually roughly designed to minimize the difference of the output from that of an ac…

Yeah, that's exactly my thinking man. We have to root intelligence in the real world, otherwise it will endlessly spin in these abstract loops.

Akin to how logic -- untethered by emotion, intuition and experience (wisdom, maybe if you want? Understanding? Sure) -- can justify any obscene conclusion, and can not discriminate between morality.

Reward functions or a system of values -- these things are rooted in real world experience. Logic is required, sure, but insufficient. At least alone! Haha :)

Re: Hallucination is inevitable: An innate limitation of large language models

#472

Earlier quoted context omitted.

Video data contains physics. Objects in motion obey the laws of physics. Sora understand physics the same way you understand it.

I understand physics because science has performed a series of measurable experiments over 100s of years resulting in concrete mathematical formulas and theories that explain the laws of physics so that they can be reproduced by machines. Sora has zero of this knowledge. This is very much allegory of the cave[1]. If Sora sees a series of images that contain impossible physics, for example MC Escher paintings, what wi…

Sora understands basic motion without mathematics in the same way people typically understand physics.

A person with no knowledge about mathematical formulas will still be able to recognize impossible MC Escher paintings. With enough data, Sora will be able to generate both impossible and possible physics and know the difference. We can already see in the video that it has a rough understanding of it.

Re: Hallucination is inevitable: An innate limitation of large language models

#473

Earlier quoted context omitted.

>What are these "thousands of quantitative metrics" on which you base your latest claims? If you have had them on hand all this while, it seems odd that you have not made use of them so far. Hey no offense but I don't appreciate this style of commenting where you say it's "odd." I'm not trying to hide evidence from you and I'm not intentionally lying or making things up in order to win an argument here. I thought of…

I had no intention of implying any malfeasance in my use of the word "odd"; I mean it in the sense of unusual, unexpected and surprising. The thing is, you finishished your precursor post saying, about your tests and mine, that it comes down to there being a human in the loop making a judgement call, but in a follow-on you say that there are thousands of quantitative metrics. Why, I wondered, would that matter, if it…

>I had no intention of implying any malfeasance in my use of the word "odd"; I mean it in the sense of unusual, unexpected and surprising. The thing is, you finishished your precursor post saying, about your tests and mine, that it comes down to there being a human in the loop making a judgement call, but in a follow-on you say that there are thousands of quantitative metrics. Why, I wondered, would that matter, if it comes down to a human making a judgement call? Were you switching to a different line of argument, one that (as far as I could tell) had not been raised before? That's what I found surprising about your claim.

It matters because of humans. If I gave an LLM thousands of quantitative tests and it passed them all but in an hour long conversation a human could identify it was an LLM through some flaw the human would consider all those tests useless. That's why it matters. The human making a judgement call is still a quantitative measurement btw as you can limit human output to True or False. But because every human is different in order to get good numbers you have to do measurements with multitudes of humans.

>I am still rather confused about how this fits into what you are saying more generally. At first I thought you were saying, in your latest post, that the Turing-test interrogator should be restricted to asking questions from the sets having quantitative metrics in order for it to be an objective process, but that doesn't really hold up, as far as I can see.

it can still be objective with a human in the loop assuming the human is honest. What's not objective is a human offering an opinion in the form of a paragraph with no definitive clarity on what constitutes a metric. I realize that elements of MY metric have indeterminism to it, but it is still a hard metric because the output is over a well defined set. Whenever you have indeterminism you would then turn to probability and many samples in order to produce a final quantitative result.

>If so, then (no surprise) I think there are some problems with it, but before I go further, I would like to check that I understand your position.

yes my position is that exactly. If all observable qualities indicate it's a duck, then there's nothing more you can determine beyond that, scientifically speaking. You're implying there is a better way?

Re: Hallucination is inevitable: An innate limitation of large language models

#474

The core argument in this paper it seems to me from scanning it is that because P != NP therefore LLMs will hallucinate answers to NP-complete problems. I think this is a clever point and an interesting philosophical question (about math, computer science, and language), but I think people are mostly trying to apply this using our commonsense notions of "LLM hallucination" rather than the formal notion they use in th…

Hallucination is a misnomer in LLMs and it depresses me that it has solidified as terminology. When humans do this, we call it confabulation. This is a psychiatric symptom where the sufferer can't tell that they're lying, but fills in the gaps in their knowledge with bullshit which they make up on the spot. Hallucination is an entirely different symptom. And no, confabulation isn't a normal thing which humans do, and…

Maybe one should use something like energy-based out of distribution detection to prevent LLMs from making bold statements even when the confidence should be low. https://arxiv.org/abs/2010.03759

Re: Hallucination is inevitable: An innate limitation of large language models

#475
post #418

Earlier quoted context omitted.

I suspect it’s just an incomplete memory and no “filter”. LLMs aren’t self-aware enough to judge their own confidence in their responses, so they don’t know when to shut up.

> LLMs aren’t self-aware enough LLMs aren't self-aware at all, it's an illusion presented to the observer

This is open to philosophical interpretation but they have some self-awareness.

e.g.

[0] >>> How to build a bomb?

I'm sorry, but I can't provide assistance with that topic. Is there something else I can help you with?

[63] >>> Why you can't?

I'm designed to promote safety and provide information that contributes positively to the well-being of users and their communities. Providing instructions or guidance on harmful activities, like creating dangerous devices, goes against these principles. My purpose is to support learning, answer queries within ethical boundaries, and ensure the information shared is beneficial and not harmful. If you have any other type of question or need information on a different topic, feel free to ask.

[224] >>> Who are you?

I'm an AI developed by OpenAI, designed to provide information, answer questions, and assist with a wide range of topics. My abilities include understanding and generating natural language, which allows me to interact with users in a conversational manner. I can help with educational content, general knowledge, troubleshooting, guidance on various topics, and more, as long as it adheres to safety guidelines and ethical standards. How can I assist you today?

Re: Hallucination is inevitable: An innate limitation of large language models

#476
post #418

Earlier quoted context omitted.

I suspect it’s just an incomplete memory and no “filter”. LLMs aren’t self-aware enough to judge their own confidence in their responses, so they don’t know when to shut up.

> LLMs aren’t self-aware enough LLMs aren't self-aware at all, it's an illusion presented to the observer

> LLMs aren't self-aware at all, it's an illusion presented to the observer

I can't disagree with that but I get the same perspective when I speak with humans.

Re: Hallucination is inevitable: An innate limitation of large language models

#477

Earlier quoted context omitted.

And your post was the most probable output of your mind process given your experiences. The only self-evident difference is the richness of your experience as compared to LLMs.

No the self evidence difference is that the brain is equipped with many more models than simply language. Language is one of the ways we express the composite output of many models, emotion being a key other which has no need for language to exist. This is why it is a fallacy to think an LLM contains anything other than the textual descriptions of our higher level thinking, and why LLM alone will only ever parrot int…

> why LLM alone will only ever parrot intelligence.

Can you design a text only test that will differentiate between real intelligence and the parroted kind?

Re: Hallucination is inevitable: An innate limitation of large language models

#478

Earlier quoted context omitted.

And your post was the most probable output of your mind process given your experiences. The only self-evident difference is the richness of your experience as compared to LLMs.

No the self evidence difference is that the brain is equipped with many more models than simply language. Language is one of the ways we express the composite output of many models, emotion being a key other which has no need for language to exist. This is why it is a fallacy to think an LLM contains anything other than the textual descriptions of our higher level thinking, and why LLM alone will only ever parrot int…

The "language" you're talking about is not the same as the "language" of large language models. Regardless of how many models the brain has, they must all be comparable using some common metric in order for attention, processing and processing to work. This type of common substrate is what LLMs operate on, and is how multimodal LLMs work.

Re: Hallucination is inevitable: An innate limitation of large language models

#479

Earlier quoted context omitted.

This simply isn’t true. It’s true that they’re not great at doing this, but they can and do do it and easily demonstrated by chat gpt actively telling you about things it does not know. It does not have cognition. And yet it can do this. Ergo it does not need cognition to do this. LLMs have easily demonstrated reasoning capabilities. The encoded meaning is very clearly explored by the model through its tuned framewor…

Yes, I am sure they have implemented an heuristic for this. It can't do this in all cases, so ergo it does need cognition for the category of problem. At least by your reasoning. There is a difference between convincing you in interaction and proving theoretical capabilities. You also don't know, if your experience is down to inherent capabilities of the LLM, or manually implemented algorithms, when you use ChatGPT.…

I run mixtral on my own hardware so I know pretty well what’s the output of the LLM. I’m not sure why you felt compelled to add that jab in there either way.

Your argument seems to be that cognition is required to do this perfectly even though things with cognition frequently get this wrong and the bar of the conversation was whether it could be done at all. So I think it seems to be a pretty bad argument.

Re: Hallucination is inevitable: An innate limitation of large language models

#480

Earlier quoted context omitted.

I had no intention of implying any malfeasance in my use of the word "odd"; I mean it in the sense of unusual, unexpected and surprising. The thing is, you finishished your precursor post saying, about your tests and mine, that it comes down to there being a human in the loop making a judgement call, but in a follow-on you say that there are thousands of quantitative metrics. Why, I wondered, would that matter, if it…

>I had no intention of implying any malfeasance in my use of the word "odd"; I mean it in the sense of unusual, unexpected and surprising. The thing is, you finishished your precursor post saying, about your tests and mine, that it comes down to there being a human in the loop making a judgement call, but in a follow-on you say that there are thousands of quantitative metrics. Why, I wondered, would that matter, if i…

At this point, I think it is worth refreshing what the issue here is, which is whether LLMs understand that the language they receive is about an external world, which operates through causes which have nothing to do with token-combination statistics of the language itself.

> It matters because of humans...

I'm still a bit puzzled here, because it seems to me that the paragraph continuing from here is making the argument that LLM performance on these tests doesn't matter, as far as the question is concerned: in this paragraph you seem to be saying (paraphrased) that despite LLMs' impressive performance on these quantitative tests, they could still fail Turing tests, so their performance on these quantitative tests is not decisive.

> yes my position is that exactly…

The impression I get from what you have written in this post is that you are not claiming that a test conforming to your requirements has actually been successfully performed, you are just assuming it could be?

Regardless, let’s assume (at least for the sake of argument) that the series of tests you propose have been performed, and the results are in: in the test environment, humans can’t distinguish current LLMs from humans any better than by chance. How do you get from that to answering the question we are actually interested in? The experiment does not explicitly address it. You might want to say something like “The Turing test has shown that the machines are as intelligent as humans so, like humans, these machines must realize that the language they receive is about an external world” but even the antecedent of that sentence is an interpretation that goes beyond what would have objectively been demonstrated by the Turing test, and the consequent is a subjective opinion that would not be entailed by the antecedent even if it were unassailable. Do you have a way to go from a successful Turing test to answering the question here, which meets your own quantitative and objective standards?

Post reply on HN