Live data from Hacker News

An LLM is a lossy encyclopedia

simonwillison.net

211–220 of 365 posts

Re: An LLM is a lossy encyclopedia

#214

Earlier quoted context omitted.

> It’s a lossy encyclopedia that can lie to and manipulate you. So can a traditional encyclopedia.

True, but in that case we call it “errors” or “propaganda”, depending on the context and source. Plus the steep costs of traditional encyclopedias, the need to refresh collections with new data periodically, and the role of librarians, all acted as a deterrent against lying (since they’re reference material). Wikipedia can also lie, obviously, but it at least requires sources to be cited, and I can dig deeper into to…

> Normalizing LLMs as “lossy encyclopedias” is a dangerous trend in my opinion, because it effectively handwaves the need for critical thinking skills associated with research and complex task execution

Calling them "lossy encyclopedias" isn't intended as a compliment! The whole point of the analogy is to emphasize that using them in place of an encyclopedia is a bad way to apply them.

Re: An LLM is a lossy encyclopedia

#215
post #135

Earlier quoted context omitted.

> the user will at least need to know something about the topic beforehand. I used ChatGPT 5 over the weekend to double check dosing guidelines for a specific medication. "Provide dosage guidelines for medication [insert here]" It spit back dosing guidelines that were an order of magnitude wrong (suggested 100mcg instead of 1mg). When I saw 100mcg, I was suspicious and said "I don't think that's right" and it quickly…

Using a LLM for medical research is just as dangerous as Googling it. Always ask your doctors!

Plot twist, your doctor is looking it up on WebMD themselves

Re: An LLM is a lossy encyclopedia

#217

Please, everybody, preserve your records. Preserve your books, preserve your downloaded files (that can't be tampered with), keep everything. AI is going to make it harder and harder to find out the truth about anything over the next few years. You have a moral duty to keep your books, and keep your locally-stored information.

To that end, it seems as though archive.org will important for an entirely new reason. Not for the loss of information, but the degradation of it.

Re: An LLM is a lossy encyclopedia

#218

Please, everybody, preserve your records. Preserve your books, preserve your downloaded files (that can't be tampered with), keep everything. AI is going to make it harder and harder to find out the truth about anything over the next few years. You have a moral duty to keep your books, and keep your locally-stored information.

I get very annoyed when llms respond with quotes around certain things I ask for, then when I say what is the source of that quote? they say oh I was paraphrasing and that isnt a real quote.

At least wikipedia has sources that probably support what it says and normally the quotes are real quotes. LLMs just seem to add quotation marks as, "proof" that its confident something is correct.

Re: An LLM is a lossy encyclopedia

#219

Earlier quoted context omitted.

Not really. The whole "inference errors will always compound" idea was popular in GPT-3.5 days, and it seems like a lot of people just never updated their knowledge since. It was quickly discovered that LLMs are capable of re-checking their own solutions if prompted - and, with the right prompts, are capable of spotting and correcting their own errors at a significantly-greater-than-chance rate. They just don't do it…

You seem to be responding to a strawman, and assuming I think something I don't think. As of today, 'bad' generations early in the sequence still do tend towards responses that are distant to the ideal response. This is testable/verifiable by pre-filling responses, which I'd advise you to experiment with for yourself. 'Bad' generations early in the output sequence are somewhat mitigatable by injecting self-reflection…

With better reasoning training, the models mitigate more and more of that entirely by themselves. They "diverge into a ditch" less, and "converge towards the right answer" more. They are able to use more and more test-time compute effectively. They bring their own supply of "wait".

OpenAI's in-house reasoning training is probably best in class, but even lesser naive implementations go a long way.

Re: An LLM is a lossy encyclopedia

#220

Earlier quoted context omitted.

I'd know cutting-edge linguistics and signaling theory well beyond Shannon to parse this, not NLP or engineering reduction. What I've stated is extremely coherent to Systemic Functional Linguists. Beyond this point engineers actually have to know what signaling is, rather than 'information.' https://www.sciencedirect.com/science/article/abs/pii/S00033... Ultimately, engineering chose the wrong approach to automating…

One of the main takeaways from The Bitter Lesson was that you should fire your linguists. GPT-2 knows more about human language than any linguist could ever hope to be able to convey. If you're hitching your wagon to human linguists, you'll always find yourself in a ditch in the end.

Sorry, 2 billion years of neurobiology beats 60 years of NLP/LLMs which knows less to nothing about language since "arbitrary points can never be refined or defined to specifics" check your corners and know your inputs.

The bill is due on NLP.

Post reply on HN