Live data from Hacker News

An LLM is a lossy encyclopedia

simonwillison.net

141–150 of 365 posts

Re: An LLM is a lossy encyclopedia

#141
post #121

Earlier quoted context omitted.

Point is that it's also exactly analogous to human intelligence. There's almost nothing else to it.

This is how you spot hype nonsense - claims that anything is analogous to human intelligence. Even absent all other objections, we don't understand the human mind well enough to make a claim like that.

You don't need to understand the human mind on a mechanistic level. You only need to examine how the whole organism learns, acts, and reacts to stimulus and situation.

Even something as simple as catching a ball is basically predictive. You predict where the ball will be along its arc when it reaches a point in space where you can catch it. Then, strictly informed by that prediction, you solve a problem of motion through space -- and some very simple-seeming problems of motion through space can't be cracked in a general case without a very powerful supercomputer -- to physically catch the ball.

That's a very simple example. The major component of what we call intelligence is purely predictive. Of course Bayesian inference also works the same way, etc.

Re: An LLM is a lossy encyclopedia

#142

It’s a lossy encyclopedia that can lie to and manipulate you. In that use case, it’s fairly useless because you cannot intrinsically trust its answers without performing additional testing and research, in which case you would’ve been better off learning new things than making sure an LLM wasn’t lying to you.

> It’s a lossy encyclopedia that can lie to and manipulate you.

So can a traditional encyclopedia.

Re: An LLM is a lossy encyclopedia

#143
post #135

I totally agree with the author. Sadly, I feel like that's not what the majority of LLM users tend to view LLMs. And it's definitely not what AI companies marketing. > The key thing is to develop an intuition for questions it can usefully answer vs questions that are at a level of detail where the lossiness matters the problem is that in order to develop an intuition for questions that LLMs can answer, the user will…

> the user will at least need to know something about the topic beforehand. I used ChatGPT 5 over the weekend to double check dosing guidelines for a specific medication. "Provide dosage guidelines for medication [insert here]" It spit back dosing guidelines that were an order of magnitude wrong (suggested 100mcg instead of 1mg). When I saw 100mcg, I was suspicious and said "I don't think that's right" and it quickly…

Using a LLM for medical research is just as dangerous as Googling it. Always ask your doctors!

Re: An LLM is a lossy encyclopedia

#144
post #42

Earlier quoted context omitted.

Completely agree with you - LLMs with access to search tools that know how to use them (o3, GPT-5, Claude 4 are particularly good at this) mostly paper over the problems caused by a lossy set of knowledge in the model weights themselves. But... end users need to understand this in order to use it effectively. They need to know if the LLM system they are talking to has access to a credible search engine and is good at…

From earlier today: Me: How do I change the language settings on YouTube? Claude: Scroll to the bottom of the page and click the language button on the footer. Me: YouTube pages scroll infinitely. Claude: Sorry! Just click on the footer without scrolling, or navigate to a page where you can scroll to the bottom like a video. (Videos pages also scroll indefinitely through comments) Me: There is no footer, you're just…

IME, eventually, after a long time, the scrolling stops and you can get to the footer. YMMV!

Re: An LLM is a lossy encyclopedia

#145
post #135

I totally agree with the author. Sadly, I feel like that's not what the majority of LLM users tend to view LLMs. And it's definitely not what AI companies marketing. > The key thing is to develop an intuition for questions it can usefully answer vs questions that are at a level of detail where the lossiness matters the problem is that in order to develop an intuition for questions that LLMs can answer, the user will…

> the user will at least need to know something about the topic beforehand. I used ChatGPT 5 over the weekend to double check dosing guidelines for a specific medication. "Provide dosage guidelines for medication [insert here]" It spit back dosing guidelines that were an order of magnitude wrong (suggested 100mcg instead of 1mg). When I saw 100mcg, I was suspicious and said "I don't think that's right" and it quickly…

Perhaps the absolute worst use-case for an LLM

Re: An LLM is a lossy encyclopedia

#147
I recently remembered a question I had a decade ago.

Why 1 nostril is always clogged up when breathing and it seems to switch now and then?

It's magnificent that I can finally get an answer for that and would never imagine it is completely natural. Never learned about it at school.

No way I can find this via search engine as it just gives me SEO garbage or anecdotal silliness.

I've been going back and getting answers to many questions I previously couldn't.

Re: An LLM is a lossy encyclopedia

#148
post #71

Every encyclopedia is lossy, by definition. Even the most expansive holds a tiny fraction of human knowledge (which is a fraction of what we could know). On the other hand, it’s not worse than other analogies.

This.

Also every encyclopedia is full of things that are wrong. People seem to be forgetting this basic issue. Any given authoritative, well respected source will contain mistakes, errors of omission, and downright lies. Part of a proper education used to be that when writing things, you need to site your sources, and the sources can't just be an encyclopedia. I use LLMs a lot, but if anything is really important, I'm going to fact check it and look for other sources.

Re: An LLM is a lossy encyclopedia

#149
I did a similar experiment and found that GPT5 hallucinates upto 20% in domains like cricket stats where there is too much info to memorize. However interestingly the mini version refuses to answer most of the time which is a better approach imho. https://kaamvaam.com/machine-learning-ai/llm-eval-hallucinat...
Post reply on HN