Live data from Hacker News

An LLM is a lossy encyclopedia

simonwillison.net

171–180 of 365 posts

Re: An LLM is a lossy encyclopedia

#171
Perhaps this also explains why I almost think LLMs are not helpful in engineering-level projects:

1. My project involved programming languages and APIs that iterated several levels faster than LLM could publish a book

2. I have lost faith in LLMs developing software.

An example is the famous Unity game engine, but LLM has not helped its Unity DOTS architecture (ECS mode compared with GameObject mode). Although I have a basic understanding of it, both the entities API documentation and the LLM answers are terrible. I chose Unity because I heard it is mature so I think LLMs would be helpful with so much materials. Sadly, for ECS it doesn't. So I chose Bevy, a game engine that I can understand and apply by reading documents and can solve problems without the help of LLM.

Re: An LLM is a lossy encyclopedia

#172

Earlier quoted context omitted.

The problem is that language doesn't produce itself. Re-checking, correcting error is not relevant. Error minimization is not the fount of survival, remaining variable for tasks is. The lossy encyclopedia is neither here nor there, it's a mistaken path: "Language, Halliday argues, "cannot be equated with 'the set of all grammatical sentences', whether that set is conceived of as finite or infinite". He rejects the us…

Sorry, what? This is borderline incoherent.

The units themselves are meaningless without context. The point of existence, action, tasks is to solve the arbitrariness in language. Tasks refute language, not the other way around. This may be incoherent as the explanation is scientific, based in the latest conceptualization of linguistics.

CS never solved the incoherence of language, conduit metaphor paradox. It's stuck behind language's bottleneck, and it do so willingly blind-eyed.

Re: An LLM is a lossy encyclopedia

#173

It’s a lossy encyclopedia that can lie to and manipulate you. In that use case, it’s fairly useless because you cannot intrinsically trust its answers without performing additional testing and research, in which case you would’ve been better off learning new things than making sure an LLM wasn’t lying to you.

> It’s a lossy encyclopedia that can lie to and manipulate you. So can a traditional encyclopedia.

True, but in that case we call it “errors” or “propaganda”, depending on the context and source. Plus the steep costs of traditional encyclopedias, the need to refresh collections with new data periodically, and the role of librarians, all acted as a deterrent against lying (since they’re reference material).

Wikipedia can also lie, obviously, but it at least requires sources to be cited, and I can dig deeper into topics at my leisure or need in order to improve my knowledge.

I cannot do either with an LLM. It is not obligated to cite sources, and even if it is it can just make shit up that’s impossible to follow or leads back to AI-generated slop - self-referencing, in other words. It also doesn’t teach you (by default, and my opinions of its teaching skills are an entirely different topic), but instead gives you an authoritative answer in tone, but not in practice.

Normalizing LLMs as “lossy encyclopedias” is a dangerous trend in my opinion, because it effectively handwaves the need for critical thinking skills associated with research and complex task execution, something in sore supply in the modern, Western world.

Re: An LLM is a lossy encyclopedia

#174

Earlier quoted context omitted.

Sorry, what? This is borderline incoherent.

The units themselves are meaningless without context. The point of existence, action, tasks is to solve the arbitrariness in language. Tasks refute language, not the other way around. This may be incoherent as the explanation is scientific, based in the latest conceptualization of linguistics. CS never solved the incoherence of language, conduit metaphor paradox. It's stuck behind language's bottleneck, and it do so…

What? This is even less coherent.

You weren't talking to GPT-4o about philosophy recently, were you?

Re: An LLM is a lossy encyclopedia

#175

Earlier quoted context omitted.

> It’s a lossy encyclopedia that can lie to and manipulate you. So can a traditional encyclopedia.

We're at such a strange point where even school children knows that something like Wikipedia isn't necessarily factually correct and that you need to double check. They then go and ask ChatGPT, as if it wasn't trained on Wikipedia. We haven't reached the stage yet where the majority of people are as sceptical of chatbots as they are of Wikipedia. I get that even if people know not to trust a wiki, they might anyway,…

To be fair, most people aren’t even critical of Wikipedia. They read an article, consume its content, and believe themselves competent experts without digging into the sources, the papers, or the talk pages for discourse and dissent.

Giving LLMs credibility as “lossless encyclopedias” is tacit approval of further dumbing-down of humanity through answer engines instead of building critical thinking skills.

Re: An LLM is a lossy encyclopedia

#176

Earlier quoted context omitted.

> the user will at least need to know something about the topic beforehand. This is why I've said a few times here on HN and elsewhere, if you're using an LLM you need to think of yourself as an architect guiding a Junior to Mid Level developer. Juniors can do amazing things, they can also goof up hard. What's really funny is you can make them audit their own code in a new context window, and give you a detailed answ…

> if you're using an LLM you need to think of yourself as an architect guiding a Junior to Mid Level developer. The thing is coding can (and should) be part of the design process. Many times, I though I have a good idea of what the solution should look like, then while coding, I got exposed more to the libraries and other parts of the code, which led me to a more refined approach. This exposure is what you will miss…

I agree. I mostly use it for scaffolding, I don't like letting it do all the work for me.

Re: An LLM is a lossy encyclopedia

#177

Earlier quoted context omitted.

The units themselves are meaningless without context. The point of existence, action, tasks is to solve the arbitrariness in language. Tasks refute language, not the other way around. This may be incoherent as the explanation is scientific, based in the latest conceptualization of linguistics. CS never solved the incoherence of language, conduit metaphor paradox. It's stuck behind language's bottleneck, and it do so…

What? This is even less coherent. You weren't talking to GPT-4o about philosophy recently, were you?

I'd know cutting-edge linguistics and signaling theory well beyond Shannon to parse this, not NLP or engineering reduction. What I've stated is extremely coherent to Systemic Functional Linguists.

Beyond this point engineers actually have to know what signaling is, rather than 'information.'

https://www.sciencedirect.com/science/article/abs/pii/S00033...

Ultimately, engineering chose the wrong approach to automating language, and it sinks the field. It's irreversible.

Re: An LLM is a lossy encyclopedia

#178

Earlier quoted context omitted.

> You never have a clear JPEG of a lamp, compress it, and get a clear image of the Milky Way, then reopen the image and get a clear image of a pile of dirt. Oh but it's much worse than that: because most LLMs aren't deterministic in the way they operate [1], you can get a pristine image of a different pile of dirt every single time you ask. [1] there are models where if you have the "model + prompt + seed" you're at…

"Deterministic" is overrated. Computers are deterministic. Most of the time. If you really don't think about all the times they aren't. But if you leave the CPU-land and go out into the real world, you don't have the privilege of working with deterministic systems at all. Engineering with LLMs is closer to "designing a robust industrial process that's going to be performed by unskilled minimum wage workers" than it i…

And one major issue is that LLMs are largely being sold and understood more like reliable algorithms than what they really are.

If everyone understood the distinction and their limitations, they wouldn’t be enjoying this level of hype, or leading to teen suicides and people giving themselves centuries-old psychiatric illnesses. If you “go out into the real world” you learn people do not understand LLMs aren’t deterministic and that they shouldn’t blindly accept their outputs.

https://archive.ph/rdL9W

https://archive.ph/20241023235325/https://www.nytimes.com/20...

https://archive.ph/20250808145022/https://www.404media.co/gu...

Re: An LLM is a lossy encyclopedia

#179
post #135

Earlier quoted context omitted.

> the user will at least need to know something about the topic beforehand. I used ChatGPT 5 over the weekend to double check dosing guidelines for a specific medication. "Provide dosage guidelines for medication [insert here]" It spit back dosing guidelines that were an order of magnitude wrong (suggested 100mcg instead of 1mg). When I saw 100mcg, I was suspicious and said "I don't think that's right" and it quickly…

What if you had told it again that you don't think that's right? Would it have stuck to it's guns and went "oh, no, I am right here" or would it have backed down and said "Oh, silly me, you're right, here's the real dosage!" and give you again something wrong? I do agree that to get the full usage out of an LLM you should have some familiarity with what you're asking about. If you didn't already have a sense of what…

I replied in the same thread "Are you sure that sounds like a low dose". It stuck to the (correct) recommendation in the 2nd response, but added in a few use cases for higher doses. So seems like it stuck to its guns for the most part.

For things like this, it would definitely be better for it to act more like a search engine and direct me to trustworthy sources for the information rather than try to provide the information directly.

Re: An LLM is a lossy encyclopedia

#180
post #164
post #117

Earlier quoted context omitted.

The argument is that a banana is a squishy hammer. You're saying hammers shouldn't be squishy. Simon is saying don't use a banana as a hammer.

> You're saying hammers shouldn't be squishy. No, that is not what I’m saying. My point is closer to “the words chosen to describe the made up concept do not translate to the idea being conveyed”. I tried to make that fit into your idea of the banana and squishy hammer, but now we’re several levels of abstraction deep using analogies to discuss analogies so it’s getting complicated to communicate clearly. > Simon is…

This is the type of comment that has been killing HN lately. “I agree with you but I want to disagree because I’m generally just that type of person. Also I am unable to tell my disagreeing point adds nothing.”
Post reply on HN