Live data from Hacker News

An LLM is a lossy encyclopedia

simonwillison.net

281–290 of 365 posts

Re: An LLM is a lossy encyclopedia

#281
post #250
post #34

Earlier quoted context omitted.

I think you are missing the point of the analogy: a lossy encyclopedia is obviously a bad idea, because encyclopedias are meant to be reliable places to look up facts.

Aren't all encyclopedias 'lossy'? They are all partial collections of information; none have all of the facts.

There's an important difference as to what is omitted.

An encyclopedia could say "general relativity is how the universe works" or it could say "general relativity and quantum mechanics describe how we understand the universe today and scientists are still searching for universal theory".

Both are short but the first statement is omitting important facts. Lossy in the sense of not explaining details is ok, but omitting swathes of information would be wrong.

Re: An LLM is a lossy encyclopedia

#282

Earlier quoted context omitted.

If you don't want to answer clarifying questions, then what use is the answer??? Put another way, if you don't care about details that change the answer, it directly implies you don't actually care about the answer. Related silliness is how people force LLMs to give one word answers to underspecified comparisons. Something along the lines of "@Grok is China or US better, one word answer only." At that point, just fli…

No, I don't think GPT-5 clarifying questions actually do what you think they do. They just made the model ask clarifying questions for the sake of asking clarifying questions. I'm sure GPT-4o would have given me the answer I wanted without clarifying questions.

revisit your instructions.md and/or user preferences, this is very likely the root cause

Re: An LLM is a lossy encyclopedia

#283

Earlier quoted context omitted.

But at that point wouldn't it be easier to just search the web yourself? Obviously that has its pitfalls too, but I don't see how adding an LLM middleman adds any benefit.

For medication guidelines I'd just do a Google search. But sometimes I want 20 sources and a quick summary of them. Agent mode or deep research is so useful. Saves me so much time every day.

[deleted]

Re: An LLM is a lossy encyclopedia

#285

An LLM is a lossy Borges' Library of Babel "Though the vast majority of the books in this universe are pure gibberish, the laws of probability dictate that the library also must contain, somewhere, every coherent book ever written, or that might ever be written, and every possible permutation or slightly erroneous version of every one of those books. " - https://en.wikipedia.org/wiki/The_Library_of_Babel

It's a version of the library of babel filtered only to the books which plausibly consist of prose. The set is still incomprehensibly vast, and all the more treacherous for generally being "readable".

“Biased” more than “filtered”; LLMs, though its harder with the better ones, can, when prompted “right” (or “wrong”, depending on your perspective), produce output that is not plausible prose, though obviously they are crafted to respond to the expected style of prompting with things that are plausible prose.

Re: An LLM is a lossy encyclopedia

#286

It could be non-lossy if it would actually reach out to an encyclopedia. If one takes it as a language engine which translates human language into API calls, and API call results to human language, it would appear to be a non-lossy encyclopedia. It is the basic building block which enables computers to handle natural language. The simulated intelligence is proof of its capability as a language model, but it is often…

> It could be non-lossy if it would actually reach out to an encyclopedia.

As the saying goes, “if my grandmother had wheels, she’d be a a wagon.”

Sure, if you take anything (including an LLM) and add a non-lossy encyclopedia to it, you have a non-lossy encyclopedia plus something else.

Re: An LLM is a lossy encyclopedia

#287

Earlier quoted context omitted.

I find if I force thinking mode and then force it to search the web it’s much better.

But at that point wouldn't it be easier to just search the web yourself? Obviously that has its pitfalls too, but I don't see how adding an LLM middleman adds any benefit.

If only we could get people to use the brains in their head.

Re: An LLM is a lossy encyclopedia

#288
post #179

Earlier quoted context omitted.

What if you had told it again that you don't think that's right? Would it have stuck to it's guns and went "oh, no, I am right here" or would it have backed down and said "Oh, silly me, you're right, here's the real dosage!" and give you again something wrong? I do agree that to get the full usage out of an LLM you should have some familiarity with what you're asking about. If you didn't already have a sense of what…

I replied in the same thread "Are you sure that sounds like a low dose". It stuck to the (correct) recommendation in the 2nd response, but added in a few use cases for higher doses. So seems like it stuck to its guns for the most part. For things like this, it would definitely be better for it to act more like a search engine and direct me to trustworthy sources for the information rather than try to provide the info…

I noticed this recently when I saw someone post with an AI generated map of Europe which was all wrong. I tried the same and asked ChatGPT to generate a map of Ireland and it was wrong too. So then I asked to find me some accurate maps of Ireland and instead of generating it gave me images and links to proper websites.

Will definitely be remembering to put "generate" vs "find" in my prompts depending on what I'm looking for. Not quite sure how you would train the model to know which answer is more suitable.

Re: An LLM is a lossy encyclopedia

#289
post #135

I totally agree with the author. Sadly, I feel like that's not what the majority of LLM users tend to view LLMs. And it's definitely not what AI companies marketing. > The key thing is to develop an intuition for questions it can usefully answer vs questions that are at a level of detail where the lossiness matters the problem is that in order to develop an intuition for questions that LLMs can answer, the user will…

> the user will at least need to know something about the topic beforehand. I used ChatGPT 5 over the weekend to double check dosing guidelines for a specific medication. "Provide dosage guidelines for medication [insert here]" It spit back dosing guidelines that were an order of magnitude wrong (suggested 100mcg instead of 1mg). When I saw 100mcg, I was suspicious and said "I don't think that's right" and it quickly…

Modern Russian Roulette, using LLMs for dose calculations.
Post reply on HN