Live data from Hacker News

An LLM is a lossy encyclopedia

simonwillison.net

191–200 of 365 posts

Re: An LLM is a lossy encyclopedia

#192
Please, everybody, preserve your records. Preserve your books, preserve your downloaded files (that can't be tampered with), keep everything. AI is going to make it harder and harder to find out the truth about anything over the next few years.

You have a moral duty to keep your books, and keep your locally-stored information.

Re: An LLM is a lossy encyclopedia

#193
post #115

Earlier quoted context omitted.

I'm frustrated by the number of times I encounter people assuming that the current model behavior is inevitable. There's been hundreds of billions of dollars spent on training LLMs to do specific things. What exactly they've been trained on matters; they could have been trained to do something else. Interacting with a base model versus an instruction tuned model will quickly show you the difference between the innate…

Some of the Anthropic guys have said that the core thing holding the models back is training, and they're confident the gains will keep coming as they figure out how to onboard more and more training data. So yeah, Claude might suck at reading and writing plumbing diagrams, but they claim the barrier is simply a function of training, not any kind of architectural limitation.

I agree with the general idea, but "sucks at reading plumbing diagrams" is the one specific example where Claude is actually choked by its unfortunate architecture.

The "naive" vision implementation for LLMs is: break the input image down into N tokens and cram those tokens into the context window. The "break the input image down" part is completely unaware of the LLM's context, and doesn't know what data would be useful to the LLM at all. Often, the vision frontend just tries to convey the general "vibes" of the image to the LLM backend, and hopes that the LLM can pick out something useful from that.

Which is "good enough" for a lot of tasks, but not all of them, not at all.

Re: An LLM is a lossy encyclopedia

#194
post #121

Earlier quoted context omitted.

This is how you spot hype nonsense - claims that anything is analogous to human intelligence. Even absent all other objections, we don't understand the human mind well enough to make a claim like that.

You don't need to understand the human mind on a mechanistic level. You only need to examine how the whole organism learns, acts, and reacts to stimulus and situation. Even something as simple as catching a ball is basically predictive. You predict where the ball will be along its arc when it reaches a point in space where you can catch it. Then, strictly informed by that prediction, you solve a problem of motion thr…

> The major component of what we call intelligence is purely predictive.

Making more unfounded, nonsensical claims does not reinforce your first unfounded, nonsensical claim.

I'm sure statisticians would love it if the human mind were an inference machine, but that doesn't make it one. Your point of view on this is faith-based.

Re: An LLM is a lossy encyclopedia

#195
post #178

Earlier quoted context omitted.

And one major issue is that LLMs are largely being sold and understood more like reliable algorithms than what they really are. If everyone understood the distinction and their limitations, they wouldn’t be enjoying this level of hype, or leading to teen suicides and people giving themselves centuries-old psychiatric illnesses. If you “go out into the real world” you learn people do not understand LLMs aren’t determi…

It's nothing new. LLMs are unreliable, but in the same ways humans are.

But LLMs output is not being treated the same as human output, and that comparison is both tired and harmful. People are routinely acting like “this is true because ChatGPT said so” while they wouldn’t do the same for any random human.

LLMs aren’t being sold as unreliable. On the contrary, they are being sold as the tool which will replace everyone and do a better job at a fraction of the piece.

Re: An LLM is a lossy encyclopedia

#197

Please, everybody, preserve your records. Preserve your books, preserve your downloaded files (that can't be tampered with), keep everything. AI is going to make it harder and harder to find out the truth about anything over the next few years. You have a moral duty to keep your books, and keep your locally-stored information.

[flagged]

Re: An LLM is a lossy encyclopedia

#198

Less than 1% of an LLM is a lossy encyclopedia. The other 99+% is all of the lossy knowledge that isn't even in encyclopedias in the first place. Including going much, much, much deeper than e.g. Wikipedia in many areas. So there it's not "lossy" -- it's effectively the opposite, i.e. "super resolution". And very, very little of what I look up using LLM's is anywhere in Wikipedia to begin with.

Outdated or terrible documentation would leads LLM giving a unexpected answer, and that would mislead me!

Re: An LLM is a lossy encyclopedia

#199
post #135

Earlier quoted context omitted.

> the user will at least need to know something about the topic beforehand. I used ChatGPT 5 over the weekend to double check dosing guidelines for a specific medication. "Provide dosage guidelines for medication [insert here]" It spit back dosing guidelines that were an order of magnitude wrong (suggested 100mcg instead of 1mg). When I saw 100mcg, I was suspicious and said "I don't think that's right" and it quickly…

I find if I force thinking mode and then force it to search the web it’s much better.

But at that point wouldn't it be easier to just search the web yourself? Obviously that has its pitfalls too, but I don't see how adding an LLM middleman adds any benefit.

Re: An LLM is a lossy encyclopedia

#200

Earlier quoted context omitted.

What? This is even less coherent. You weren't talking to GPT-4o about philosophy recently, were you?

I'd know cutting-edge linguistics and signaling theory well beyond Shannon to parse this, not NLP or engineering reduction. What I've stated is extremely coherent to Systemic Functional Linguists. Beyond this point engineers actually have to know what signaling is, rather than 'information.' https://www.sciencedirect.com/science/article/abs/pii/S00033... Ultimately, engineering chose the wrong approach to automating…

One of the main takeaways from The Bitter Lesson was that you should fire your linguists. GPT-2 knows more about human language than any linguist could ever hope to be able to convey.

If you're hitching your wagon to human linguists, you'll always find yourself in a ditch in the end.

Post reply on HN