Live data from Hacker News

Overcoming the limits of current LLMs

seanpedersen.github.io

71–80 of 111 posts

Re: Overcoming the limits of current LLMs

#71

Man it seems like the ship has sailed on "hallucination" but it's such a terrible name for the phenomenon we see. It is a major mistake to imply the issue is with perception rather than structural incompetence. Why not just say "incoherent output"? It's actually descriptive and doesn't require bastardizing a word we already find meaningful to mean something completely different.

"Hallucinations" implies that someone isn't of sound mental state. We can argue forever about what that means for a LLM and whether that's appropriate, but I think it's absolutely the right attitude and approach to be taking toward these things. They simply do not behave like humans of sound minds, and "hallucinations" conveys that in a way that "confabulations" or even "bullshit" does not. (Though "bullshit" isn't b…

Bullshit is the most descriptive one.

LLMs don't do it because they are out of their right mind. They do it because every single answer they say is invented caring only about form, and not correctness.

But yeah, that ship has already sailed.

Re: Overcoming the limits of current LLMs

#72
post #32
post #28

Earlier quoted context omitted.

I disagree with this take because LLMs are, always, hallucinating. When they get things right it’s because they are lucky. Yes, yes, it’s more complicated than that, but the essence of LLMs is that they are very good at being lucky. So good that they will often give you better results than random search engine clicks, but not good enough to be useful for anything important. I think what calling the times they get thi…

To offer a satirical analogy: "Lastly, I want to reassure investors and members of the press that we take these concerns very seriously: Hindenburg 2 will only contain only normal and unreactive hydrogen gas, and not the rare and unusual explosive kind, which is merely a temporary hurdle in this highly dynamic and growing field." Edit: It retrospect, perhaps a better analogy would involve gasoline, as its explosive n…

Yes (to the edit), an analogy with making planes safer by only using non-flammable fuels is perfect.

Re: Overcoming the limits of current LLMs

#73

LLMs don't only hallucinate because of mistaken statements in their training data. It just comes hand-in-hand with the model's ability to remix, interpolate, and extrapolate answers to other questions that aren't directly answered in the dataset. For example if I ask ChatGPT a legal question, it might cite as precedent a case that doesn't exist at all (but which seems plausible, being interpolated from cases that do…

These problems are solvable, but are we really making a better world by solving them? When you ask yourself that question -- and you do ask yourself that, right? -- what's your answer?

I do, all the time! My answer is "most likely not". (I assumed that answer was implied by my expressing sadness about all the work being invested in them.) This is why, although I try to keep up-to-date with and understand these technologies, I am not being paid to develop them.

Re: Overcoming the limits of current LLMs

#74

Man it seems like the ship has sailed on "hallucination" but it's such a terrible name for the phenomenon we see. It is a major mistake to imply the issue is with perception rather than structural incompetence. Why not just say "incoherent output"? It's actually descriptive and doesn't require bastardizing a word we already find meaningful to mean something completely different.

I think it's a pretty good name for the phenomenon -- maybe the only problem with the term is that what models are doing is 100% hallucination all the time -- it's just that when the hallucinations are useful we don't call them hallucinations -- so maybe that is a problem with the term (not sure if that's what you are getting at).

But there's nothing at all different about what the model is doing between these cases -- the models are hallucinating all the time and have no ability to assess when they are hallucinating "right" or "wrong" or useful/non-useful output in any meaningful way.

Re: Overcoming the limits of current LLMs

#75

LLMs don't only hallucinate because of mistaken statements in their training data. It just comes hand-in-hand with the model's ability to remix, interpolate, and extrapolate answers to other questions that aren't directly answered in the dataset. For example if I ask ChatGPT a legal question, it might cite as precedent a case that doesn't exist at all (but which seems plausible, being interpolated from cases that do…

100% I think the author is really misunderstanding the issue here. "Hallucination" is a fundamental aspect of the design of Large Language Models. Narrowing the distribution of the training data will reduce the LLM's ability to generalize, but it won't stop hallucinations.

Re: Overcoming the limits of current LLMs

#76
post #39
post #31

We can't develop a universally coherent data set because what we understand as "truth" is so intensely contextual that we can't hope to cover the amount of context needed to make the things work how we want, not to mention the numerous social situations where writing factual statements would be awkward or disastrous. Here are a few examples of statements that are not "factual" in the sense of being derivable from a u…

> We can't develop a universally coherent data set because Yet every child seems to manage, when raised by a small village, over a period of about 18 years. I guess we just need to give these LLMs a little more love and attention.

And then you go out into the real world, talk to real adults, and discover that the majority of people don't have a coherent mental model of the world, and have completely ridiculous ideas that aren't anywhere near an approximation of the real physical world.

Re: Overcoming the limits of current LLMs

#77
post #68

Man it seems like the ship has sailed on "hallucination" but it's such a terrible name for the phenomenon we see. It is a major mistake to imply the issue is with perception rather than structural incompetence. Why not just say "incoherent output"? It's actually descriptive and doesn't require bastardizing a word we already find meaningful to mean something completely different.

I think calling it hallucination is because of our tendency to anthropomorphize things. Humans hallucinate. Programs have bugs.

The point is that this isn't a bug.

It's inherent to how LLMs work and is expected although undesired behaviour.

Re: Overcoming the limits of current LLMs

#78
post #75

LLMs don't only hallucinate because of mistaken statements in their training data. It just comes hand-in-hand with the model's ability to remix, interpolate, and extrapolate answers to other questions that aren't directly answered in the dataset. For example if I ask ChatGPT a legal question, it might cite as precedent a case that doesn't exist at all (but which seems plausible, being interpolated from cases that do…

100% I think the author is really misunderstanding the issue here. "Hallucination" is a fundamental aspect of the design of Large Language Models. Narrowing the distribution of the training data will reduce the LLM's ability to generalize, but it won't stop hallucinations.

I agree in that a perfectly consistent dataset won't completely stop statistical language models from hallucinating but it will reduce it. I think it is established that data quality is more important than quantity. Bullshit in -> bullshit out, so a focus on data quality is good and needed IMO.

I am also saying LMs output should cite sources and give confidence scores (which reflects how much the output is in or out of the training distrtibution).

Re: Overcoming the limits of current LLMs

#79

Man it seems like the ship has sailed on "hallucination" but it's such a terrible name for the phenomenon we see. It is a major mistake to imply the issue is with perception rather than structural incompetence. Why not just say "incoherent output"? It's actually descriptive and doesn't require bastardizing a word we already find meaningful to mean something completely different.

I prefer “confabulate” to describe this phenomena.

: to fill in gaps in memory by fabrication

> In psychology, confabulation is a memory error consisting of the production of fabricated, distorted, or misinterpreted memories about oneself or the world.

It’s more about coming up with a plausible explanation in the absence of a readily-available one.

Re: Overcoming the limits of current LLMs

#80

Man it seems like the ship has sailed on "hallucination" but it's such a terrible name for the phenomenon we see. It is a major mistake to imply the issue is with perception rather than structural incompetence. Why not just say "incoherent output"? It's actually descriptive and doesn't require bastardizing a word we already find meaningful to mean something completely different.

"Hallucinations" implies that someone isn't of sound mental state. We can argue forever about what that means for a LLM and whether that's appropriate, but I think it's absolutely the right attitude and approach to be taking toward these things. They simply do not behave like humans of sound minds, and "hallucinations" conveys that in a way that "confabulations" or even "bullshit" does not. (Though "bullshit" isn't b…

"Sound" minds for humans is graded on a curve, and this trick is not acknowledged, or popular.
Post reply on HN