Live data from Hacker News

An LLM is a lossy encyclopedia

simonwillison.net

201–210 of 365 posts

Re: An LLM is a lossy encyclopedia

#201
post #195

Earlier quoted context omitted.

It's nothing new. LLMs are unreliable, but in the same ways humans are.

But LLMs output is not being treated the same as human output, and that comparison is both tired and harmful. People are routinely acting like “this is true because ChatGPT said so” while they wouldn’t do the same for any random human. LLMs aren’t being sold as unreliable. On the contrary, they are being sold as the tool which will replace everyone and do a better job at a fraction of the piece.

That comparison is more useful than the alternatives. Anthropomorphic framing is one of the best framings we have for understanding what properties LLMs have.

"LLM is like an overconfident human" certainly beats both "LLM is like a computer program" and "LLM is like a machine god". It's not perfect, but it's the best fit at 2 words or less.

Re: An LLM is a lossy encyclopedia

#202
post #88

Earlier quoted context omitted.

Lossy compression does make things up. We call them compression artefacts. In compressed audio these can be things like clicks and boings and echoes and pre-echoes. In compressed images they can be ripply effects near edges, banding in smoothly varying regions, but there are also things like https://www.dkriesel.com/en/blog/2013/0802_xerox-workcentres... where one digit is replaced with a nice clean version of a diff…

> Lossy compression does make things up. We call them compression artefacts. I don’t think this is a great analogy. Lossy compression of images or signals tends to throw out information based on how humans perceive it, focusing on the most important perceptual parts and discarding the less important parts. For example, JPEG essentially removes high frequency components from an image because more information is presen…

LLM confabulations might as well be gradual in the latent space. I don’t think lossy is synonymous to perceptual and the high frequency components rather easily translate to less popular data.

Re: An LLM is a lossy encyclopedia

#204

Less than 1% of an LLM is a lossy encyclopedia. The other 99+% is all of the lossy knowledge that isn't even in encyclopedias in the first place. Including going much, much, much deeper than e.g. Wikipedia in many areas. So there it's not "lossy" -- it's effectively the opposite, i.e. "super resolution". And very, very little of what I look up using LLM's is anywhere in Wikipedia to begin with.

Outdated or terrible documentation would leads LLM giving a unexpected answer, and that would mislead me!

You think there aren't outdated or terrible articles on Wikipedia? Or the internet in general?

The internet misleads you. Which is why we develop good BS detectors, and double-check information as needed. There isn't perfect information anywhere. Even official docs are often riddled with errors and inconsistencies.

Re: An LLM is a lossy encyclopedia

#205

Earlier quoted context omitted.

Using a LLM for medical research is just as dangerous as Googling it. Always ask your doctors!

I disagree. I'd wager that state of the art LLMs can beat out of the average doctor at diagnosis given a detailed list of symptoms, especially for conditions the doctor doesn't see on a regular basis.

"Given a detailed list of symptoms" is sure holding a lot of weight in that statement. There's way too much information that doctors tacitly understand from interactions with patients that you really cannot rely on those patients supplying in a "detailed list". Could it diagnose correctly, some of the time? Sure. But the false positive rate would be huge given LLMs suggestible nature. See the half dozen news stories covering AI induced psychosis for reference.

Regardless, it's diagnostic capability is distinct from the dangers it presents, which is what the parent comment was mentioning.

Re: An LLM is a lossy encyclopedia

#207

I totally agree with the author. Sadly, I feel like that's not what the majority of LLM users tend to view LLMs. And it's definitely not what AI companies marketing. > The key thing is to develop an intuition for questions it can usefully answer vs questions that are at a level of detail where the lossiness matters the problem is that in order to develop an intuition for questions that LLMs can answer, the user will…

> The key thing is to develop an intuition for questions it can usefully answer vs questions that are at a level of detail where the lossiness matters

It's also useful to have an intuition for what things an LLM is liable to get wrong/hallucinate, one of which is questions where the question itself suggests one or more obvious answers (which may or may not be correct), which the LLM may well then hallucinate, and sound reasonable, if it doesn't "know".

Re: An LLM is a lossy encyclopedia

#208
post #30

A lossy encyclopaedia should be missing information and be obvious about it, not making it up without your knowledge and changing the answer every time . When you have a lossy piece of media, such as a compressed sound or image file, you can always see the resemblance to the original and note the degradation as it happens. You never have a clear JPEG of a lamp, compress it, and get a clear image of the Milky Way, the…

Yeah an LLM is an unreliable librarian, if anything.

Re: An LLM is a lossy encyclopedia

#209
An LLM is a lossy Borges' Library of Babel

"Though the vast majority of the books in this universe are pure gibberish, the laws of probability dictate that the library also must contain, somewhere, every coherent book ever written, or that might ever be written, and every possible permutation or slightly erroneous version of every one of those books. " -https://en.wikipedia.org/wiki/The_Library_of_Babel

Re: An LLM is a lossy encyclopedia

#210

I totally agree with the author. Sadly, I feel like that's not what the majority of LLM users tend to view LLMs. And it's definitely not what AI companies marketing. > The key thing is to develop an intuition for questions it can usefully answer vs questions that are at a level of detail where the lossiness matters the problem is that in order to develop an intuition for questions that LLMs can answer, the user will…

> The key thing is to develop an intuition for questions it can usefully answer vs questions that are at a level of detail where the lossiness matters It's also useful to have an intuition for what things an LLM is liable to get wrong/hallucinate, one of which is questions where the question itself suggests one or more obvious answers (which may or may not be correct), which the LLM may well then hallucinate, and sou…

LLMs are very sensitive to leading questions. A small hint of that the expected answer looks like will tend to produce exactly that answer.
Post reply on HN