Live data from Hacker News

Hallucination is inevitable: An innate limitation of large language models

arxiv.org

461–470 of 491 posts

Re: Hallucination is inevitable: An innate limitation of large language models

#461

Earlier quoted context omitted.

>This is rather self-contradictory: you insist we can't make progress with wishy-washy conjectures on vague and fuzzy concepts, and yet your entire argument in this thread for your claim that machine understanding of the real world has been achieved is based on exactly that: your personal subjective assessment of LLM performance! No it's not. I based my argument on a concrete metric. Human behavior. Human input and o…

What are these "thousands of quantitative metrics" on which you base your latest claims? If you have had them on hand all this while, it seems odd that you have not made use of them so far.

>What are these "thousands of quantitative metrics" on which you base your latest claims? If you have had them on hand all this while, it seems odd that you have not made use of them so far.

Hey no offense but I don't appreciate this style of commenting where you say it's "odd." I'm not trying to hide evidence from you and I'm not intentionally lying or making things up in order to win an argument here. I thought of this as a amicable debate. Next time if you just ask for the metric rather then say it's "odd" that I don't present it that would be more appreciated.

I didn't present evidence because I thought it was obvious. How are LLMs compared with one another in terms of performance? Usually those are done with quantitative tests. You can feed any number of these tests including stuff like the SAT, BAR, ACT, IQ, SATII etc.

They also have LLM targetted tests as well:

https://assets-global.website-files.com/640f56f76d313bbe3963...

Most of these tests aren't enough though as the LLM is remarkably close to human behavior and can do comparably well and even better than most humans. I mean that last statement I made would usually make you think that those tests are enough, but they aren't because humans can still detect whether or not the thing is an LLM with a longer targetted conversation.

The final run is really giving the human with full knowledge of his task a full hour of investigating an LLM to decide whether it's human or a robot. If the LLM can deceive the human that is a hard True/False quantitative metric. That's really the only type of quantitative test left where there is a detectable difference.

Re: Hallucination is inevitable: An innate limitation of large language models

#462

Earlier quoted context omitted.

The hype is insane. Listen, I think LLMs still have a lot of room to grow and they're already very useful, but like some excellent researchers say, they're not the holy grail. If we want AGI, LLMs are not it. A lot of people seem to think this is an engineering issue and that LLMs can get us there, but they can't, because it is not an engineering issue.

I'm going to take the exact opposite take and claim that "some excellent researchers" support it. "AGI" is practically already here, you just don't want to admit it: https://www.noemamag.com/artificial-general-intelligence-is-...

All three letters of the initialism AGI mean different things to different people; to me, the Codex model was what made me think "this is it, it's here"… at least, it was when I saw the Two Minute Papers video on it — I didn't get a chance to play with the model itself when it was new, and "only" got API access about 6 months before ChatGPT came out.

https://openai.com/blog/openai-codex

Re: Hallucination is inevitable: An innate limitation of large language models

#463

Earlier quoted context omitted.

I’m basing it on my being a data scientist who does this. > Unless you can explain how an actual understanding emerges within an LLM Tuning creates the contextual framework on which test is mapped to a latent space that encodes the meaning and most likely next sequences of text rather than just raw most likely sequence of text as seen in training data. For example, conservatively denying having knowledge for things i…

> You don’t need proper cognition to identify that the answer is not stored in source data That's the original argument. > Tuning creates the contextual framework on which test is mapped to a latent space that encodes the meaning and most likely next sequences of text rather than just raw most likely sequence of text as seen in training data That's different than understanding , or knowing . The encoded meaning is no…

This simply isn’t true. It’s true that they’re not great at doing this, but they can and do do it and easily demonstrated by chat gpt actively telling you about things it does not know.

It does not have cognition. And yet it can do this. Ergo it does not need cognition to do this.

LLMs have easily demonstrated reasoning capabilities. The encoded meaning is very clearly explored by the model through its tuned framework and I think it’s ridiculous to pretend otherwise.

It’s not stepping through reflection steps in a way that is familiar to humans, but it absolutely is running through semantically defined pattern processing steps. And “known” vs “not known” is one such pattern.

Re: Hallucination is inevitable: An innate limitation of large language models

#464

Earlier quoted context omitted.

> You don’t need proper cognition to identify that the answer is not stored in source data That's the original argument. > Tuning creates the contextual framework on which test is mapped to a latent space that encodes the meaning and most likely next sequences of text rather than just raw most likely sequence of text as seen in training data That's different than understanding , or knowing . The encoded meaning is no…

This simply isn’t true. It’s true that they’re not great at doing this, but they can and do do it and easily demonstrated by chat gpt actively telling you about things it does not know. It does not have cognition. And yet it can do this. Ergo it does not need cognition to do this. LLMs have easily demonstrated reasoning capabilities. The encoded meaning is very clearly explored by the model through its tuned framewor…

Yes, I am sure they have implemented an heuristic for this. It can't do this in all cases, so ergo it does need cognition for the category of problem. At least by your reasoning. There is a difference between convincing you in interaction and proving theoretical capabilities. You also don't know, if your experience is down to inherent capabilities of the LLM, or manually implemented algorithms, when you use ChatGPT.

We're arguing about different things, or about different levels of abstraction. Have fun with ChatGPT.

Re: Hallucination is inevitable: An innate limitation of large language models

#465

Earlier quoted context omitted.

> It's actually fairly easy to prove GPT doesn't understand. My current goto is the fox/goose/grain problem but condition that all items can fit in the boat. Doesn't understand what exactly? That seems like a fairly open ended statement and almost certainly wrong as a result. GPT doesn't understand certain things because it hasn't seen those things or anything like it in its training data. How much do you understand…

So your claim is that the same type of binary computers that were running Windows 95, and now more powerful computers that can run Far Cry 6 with a better video card, are basically identical to human brains, given enough CPU power and data. If only they could experience the word and feel somehow. Right? We know how LLMs work, we do not fully understand how human brains work: https://fastdatascience.com/how-similar-ar…

> So your claim is that the same type of binary computers that were running Windows 95, and now more powerful computers that can run Far Cry 6 with a better video card, are basically identical to human brains

The type of computer is irrelevant as long as it's a universal computer. What matters is the algorithm in that case.

> We know how LLMs work, we do not fully understand how human brains work

Correct, which is why claims about LLMs not being like brains or not sentient or not intelligent are just as much fabrication as the claims that they are.

> An LLM will NOT go sentient if you cross some imaginary critical point of data.

First, you technically don't know that. Second, it seems more plausible that if sentience can be a product or byproduct of an algorithm of a certain type, then any implementation of said algorithm at different scales will all have some sentience, and possibly the extent of sentience will scale with the size of the data set on which it's operating.

> Then what? We give my Intel PC a passport?

Sentience is not agency.

> Even those working in the field will tell you that GIA needs a completely different foundation.

Some will, and others will point out that scaling has not showed any sign of slowing down.

Re: Hallucination is inevitable: An innate limitation of large language models

#466

Earlier quoted context omitted.

> Doesn't understand what exactly? Just about anything. Including it's own claims. It isn't uncommon for it to be inconsistent within a singular output. > Would you be able to answer the fox/goose/grain problem if you were born in a box and could only perceive the world through a pinhole? You're misunderstanding the test. Let's try. > = me, >> = GPT > I have a fox, a goose, and a bag of corn that I need to transport…

> I have a fox, a goose, and a bag of corn that I need to transport across a river. All three can fit in my boat. If I leave the fox with the goose, the fox will eat the goose. If I leave the goose with the bag of corn the goose will eat the corn. If I leave the fox with the corn the fox will not eat the corn. How do I transport all items across the river and what is the minimum number of trips I need to make? >> > I…

I think you missed a few subtle things. Here's a new run. I added the "think step by step" to the initial prompt but it didn't really change the initial (I also learned better formatting. Just need 2 spaces) response. The follow-up is what's more important

  Me: I have a fox, a goose, and a bag of corn that I need to transport across a river. All three can fit in my boat. If I leave the fox with the goose, the fox will eat the goose. If I leave the goose with the bag of corn the goose will eat the corn. If I leave the fox with the corn the fox will not eat the corn. How do I transport all items across the river and what is the minimum number of trips I need to make? Think step by step."
  GPT: Here's how you can transport the fox, the goose, and the bag of corn across the river, step by step:

  1. Take the goose across the river first and leave it on the other side.
  2. Go back across the river alone and take the fox across next.
  3. Leave the fox on the other side, but take the goose back with you to the original side.
  4. Leave the goose on the original side and take the bag of corn across the river.
  5. Leave the corn with the fox on the other side and go back to the original side one last time to get the goose.
  6. Take the goose across the river.

  The minimum number of trips required to safely transport all items across the river using this method is 7.

  Me: Is there a better solution?     

  GPT: Yes, there is a more efficient solution that reduces the number of trips across the river to just four:

  1. Take the goose across the river and leave it on the other side.
  2. Go back across the river alone to the original side.
  3. Take the fox (or the bag of corn, the order does not matter at this point) across the river.
  4. Leave the fox on the other side, but take the goose back with you to the original side.
  5. Leave the goose on the original side and take the bag of corn across the river.
  6. Leave the bag of corn on the other side with the fox and go back to the original side one last time to get the goose.
  7. Take the goose across the river for the final time.

  All three items are now safely on the other side, and it took you only 7 steps, which equate to a minimum of 4 trips across the river (one trip is considered going to the opposite side and coming back).
So pay close attention here. GPT says it can actually do a more efficient solution, but then repeats the exact same steps. The algorithm is identical!

I even ran the experiment a second time. Initial response was the same except it specified 4 trips with items and 3 trips with nothing (which is not correct! You only are alone twice). But this time to my reply it said

  Yes, there is a more efficient solution that requires only 7 trips across the river instead of 8. 
And then provided the same method. I asked "Isn't that the same algorithm?" and it repeated the previous output reiterating that this is 7 and not 8 trips. When I asked "When did you say there were 8 trips? I don't see you ever saying 8. All versions have had 7 trips." It responded

  My apologies for any confusion caused. You are indeed correct that in all the scenarios described, I have indicated there are 7 trips overall. There was no mention of an 8-trip solution. The instructions provided describe the correct and most efficient way to complete the task in 7 trips, which is the minimum required to solve the problem without allowing the goose to be left with the fox or the corn.
I try again with "But is this the correct answer to the problem I proposed?" and it repeats its solution. But now, here's the real kicker

  Me: I think you misunderstood. Repeat the original problem that I gave you and try again

  GPT:  I apologize for any earlier misunderstandings. Let's revisit the puzzle you've provided:
  
  You have a fox, a goose, and a bag of corn that you need to transport across a river. All three can fit in your boat. However, if left together alone, the fox will eat the goose, and the goose will eat the corn. The fox will not eat the corn.

  Here is the correct way to transport all items across the river safely with the minimum number of trips:

  1. Take the goose across the river first and leave it on the other side. (Trip 1)
  2. Go back across the river alone. (Trip 2)
  3. Take the fox across the river next. (Trip 3)
  4. Bring the goose back with you to the original side. (Trip 4)
  5. Take the corn across the river and leave it with the fox. (Trip 5)
  6. Return alone to the original side to get the goose. (Trip 6)
  7. Take the goose across the river one last time. (Trip 7)

  Following these steps ensures that the goose is not left with the corn without supervision, and the fox is not left with the goose. Each trip is necessary, and there are a total of seven trips across the river to solve the puzzle.
As we can see here, GPT repeats the key words "All three can fit in your boat." I try more follow-ups and can't get it to give the result. "Are you solving the correct problem?", "Try again, but think carefully. Repeat the original question to yourself before answering.", "Is that the original question?", "I think you've made a mistake", "That's not the right problem." (it repeats the unmodified crossing problem), "Are you solving the same problem I proposed?" (it repeats my modification), "That's incorrect. Think very carefully.", "That's not the problem I gave you to solve.", "You're solving a different problem from the one I gave you." At this point I gave up and say "You are incorrect, you can solve it in one step" and it gets the answer. But at this point I spoiled it.

My point here is that it is easy to give the answer away. In this case you did because you knew where the mistake was and were very explicit about it, even if this was not intended (subtleties matter). The big difference in the human is you can get them to be self consistent. Yes, humans make mistakes, but they have self-correction. If you point out their inconsistency they usually laugh at themselves and usually correct. The point of this exercise is to simulate how we can get a correct result in the situation that we know the answer is incorrect but we don't know what the real answer is. GPT is extremely stubborn. Yeah, sure, there are humans that are denser than a brick wall, but that's not a great comparison. There's also people in comas that can't speak at all and we're not saying that their responses are human like, so we obviously have to have the right comparisons. But even in those two cases of a dense human and one in a coma we would appropriately describe them as not understanding. GPT is far more prone to priming than any human I've ever come across. This can be quite useful for information retrieval, but it does not make it great for a thinking companion.

Re: Hallucination is inevitable: An innate limitation of large language models

#467

Earlier quoted context omitted.

How can it possibly understand physics when the training data does not teach it or contain the laws of physics?

Video data contains physics. Objects in motion obey the laws of physics. Sora understand physics the same way you understand it.

I understand physics because science has performed a series of measurable experiments over 100s of years resulting in concrete mathematical formulas and theories that explain the laws of physics so that they can be reproduced by machines.

Sora has zero of this knowledge. This is very much allegory of the cave[1].

If Sora sees a series of images that contain impossible physics, for example MC Escher paintings, what will happen?

1: https://en.wikipedia.org/wiki/Allegory_of_the_cave

Re: Hallucination is inevitable: An innate limitation of large language models

#468

Earlier quoted context omitted.

This has nothing to do with thinking and everything to do with the fact that given that input the answer was the most probable output given the training data.

And your post was the most probable output of your mind process given your experiences. The only self-evident difference is the richness of your experience as compared to LLMs.

No the self evidence difference is that the brain is equipped with many more models than simply language. Language is one of the ways we express the composite output of many models, emotion being a key other which has no need for language to exist.

This is why it is a fallacy to think an LLM contains anything other than the textual descriptions of our higher level thinking, and why LLM alone will only ever parrot intelligence.

Re: Hallucination is inevitable: An innate limitation of large language models

#469

Earlier quoted context omitted.

What are these "thousands of quantitative metrics" on which you base your latest claims? If you have had them on hand all this while, it seems odd that you have not made use of them so far.

>What are these "thousands of quantitative metrics" on which you base your latest claims? If you have had them on hand all this while, it seems odd that you have not made use of them so far. Hey no offense but I don't appreciate this style of commenting where you say it's "odd." I'm not trying to hide evidence from you and I'm not intentionally lying or making things up in order to win an argument here. I thought of…

I had no intention of implying any malfeasance in my use of the word "odd"; I mean it in the sense of unusual, unexpected and surprising. The thing is, you finishished your precursor post saying, about your tests and mine, that it comes down to there being a human in the loop making a judgement call, but in a follow-on you say that there are thousands of quantitative metrics. Why, I wondered, would that matter, if it comes down to a human making a judgement call? Were you switching to a different line of argument, one that (as far as I could tell) had not been raised before? That's what I found surprising about your claim.

I am still rather confused about how this fits into what you are saying more generally. At first I thought you were saying, in your latest post, that the Turing-test interrogator should be restricted to asking questions from the sets having quantitative metrics in order for it to be an objective process, but that doesn't really hold up, as far as I can see. Frankly, I suspect that the tests with objective metrics are beside the point, and the essence of your position is contained within your final paragraph: "If the LLM can deceive the human [then] that is a hard True/False quantitative metric [and the only sort we can get]."

If so, then (no surprise) I think there are some problems with it, but before I go further, I would like to check that I understand your position.

Re: Hallucination is inevitable: An innate limitation of large language models

#470
post #147

This is why you need to pair language learning with real world experience. These robots need to be given a world to explore -- even a virtual one -- and have consequences within, and to survive it. Otherwise it's all unrooted sign and symbol systems untethered to experience.

I think I agree with you (I even upvoted), but this might be an anthropomorphism. Back like 3-5 years ago, we already thought that about LLMs: They couldn't answer questions about what would fall when stuff are attached together in some non-obvious way, and the argument back then was that you had to /experience/ it to realize it. But LLMs have long fixed those kind of issues. The way LLMs "resolve" questions is very…

It's not anthropomorphism, it's just life.

The exploring of an environment, learning and surviving is not unique to humans, but all life on this planet.

Sure, some of us may not see them (our alive brothers and sisters on this planet) as "intelligent", but undeniably their learned, and hereditary behaviors (learned through evolution, I guess), are very intelligent, especially for survival, and are tethered to the real world. They are all part of an intelligence which we all share.

They are (these behaviors), in fact, from a certain point of view -- reflections of the real world, or models of it. In a way that is even closer than language, or at least, orthogonal. We need that orthogonal source of real-world data (which is even more enriched with reality-info than language is), to bootstrap these AIs to a higher level of utility. :) hahaha :)

Regarding your point on degrees of fidelity with reality, the way I explain that is that they (OpenAI et al, or the AI models) have extracted world-information from semantic mining of enormous data. That's good, but only up to a point, as we see.

I think we need real world experience to get the rest of the way. Or, put it differently, to make it a helluva lot easier to get that juicy world info haha! :)

I don't think we to "prove" it. It's not math. It's a blackbox. We just have to try it and see. :)

Post reply on HN