Live data from Hacker News

An LLM is a lossy encyclopedia

simonwillison.net

341–350 of 365 posts

Re: An LLM is a lossy encyclopedia

#341
post #339
post #338

Earlier quoted context omitted.

Current frontier LLMs - Claude 4, GPT-5, Gemini 2.5 - are massively more likely to say "I don't know" than last year's models.

I don’t think I’ve ever seen ChatGPT 5 refuse to answer any prompt I’ve ever given it. I’m doing 20+ chats a day. What’s an example prompt where it will say “idk”? Edit: Just tried a silly one, asking it to tell me about the 8th continent on earth, which doesn’t exist. How difficult is it for the model to just say “sorry, there are only 7 continents”. I think we should expect more from LLMs and stop blaming things on…

https://chatgpt.com/share/68b85035-62ec-8006-ab20-af5931808b... - "There are only seven recognized continents on Earth: Africa, Antarctica, Asia, Australia, Europe, North America, and South America."

Here's a recent example of it saying "I don't know" - I asked it to figure out why there was an octopus in a mural about mushrooms: https://chatgpt.com/share/68b8507f-cc90-8006-b9d1-c06a227850... - "I wasn’t able to locate a publicly documented explanation of why Jo Brown (Bernoid) chose to include an octopus amid a mushroom-themed mural."

Re: An LLM is a lossy encyclopedia

#342
post #341
post #339

Earlier quoted context omitted.

I don’t think I’ve ever seen ChatGPT 5 refuse to answer any prompt I’ve ever given it. I’m doing 20+ chats a day. What’s an example prompt where it will say “idk”? Edit: Just tried a silly one, asking it to tell me about the 8th continent on earth, which doesn’t exist. How difficult is it for the model to just say “sorry, there are only 7 continents”. I think we should expect more from LLMs and stop blaming things on…

https://chatgpt.com/share/68b85035-62ec-8006-ab20-af5931808b... - "There are only seven recognized continents on Earth: Africa, Antarctica, Asia, Australia, Europe, North America, and South America." Here's a recent example of it saying "I don't know" - I asked it to figure out why there was an octopus in a mural about mushrooms: https://chatgpt.com/share/68b8507f-cc90-8006-b9d1-c06a227850... - "I wasn’t able to loca…

Not sure what your system prompt is, but asking the exact same prompt word for word for me results in a response talking about "Zealandia, a continent that is 93% submerged underwater."

The 2nd example isn't all that impressive since you're asking it to provide you something very specific. It succeeded in not hallucinating. It didn't succeed at saying "I'm not sure" in the face of ambiguity.

I want the LLM to respond more like a librarian: When they know something for sure, they tell you definitively, otherwise they say "I'm not entirely sure, but I can point you to where you need to look to get the information you need."

Re: An LLM is a lossy encyclopedia

#343
post #342
post #341

Earlier quoted context omitted.

https://chatgpt.com/share/68b85035-62ec-8006-ab20-af5931808b... - "There are only seven recognized continents on Earth: Africa, Antarctica, Asia, Australia, Europe, North America, and South America." Here's a recent example of it saying "I don't know" - I asked it to figure out why there was an octopus in a mural about mushrooms: https://chatgpt.com/share/68b8507f-cc90-8006-b9d1-c06a227850... - "I wasn’t able to loca…

Not sure what your system prompt is, but asking the exact same prompt word for word for me results in a response talking about "Zealandia, a continent that is 93% submerged underwater." The 2nd example isn't all that impressive since you're asking it to provide you something very specific. It succeeded in not hallucinating. It didn't succeed at saying "I'm not sure" in the face of ambiguity. I want the LLM to respond…

I'm using regular GPT-5, no custom instructions and memory turned off.

Can you link to your shared Zealandia result?

I think that mural result was spectacularly impressive, given that it started with a photo I took of the mural with almost no additional context.

Re: An LLM is a lossy encyclopedia

#344
post #343
post #342

Earlier quoted context omitted.

Not sure what your system prompt is, but asking the exact same prompt word for word for me results in a response talking about "Zealandia, a continent that is 93% submerged underwater." The 2nd example isn't all that impressive since you're asking it to provide you something very specific. It succeeded in not hallucinating. It didn't succeed at saying "I'm not sure" in the face of ambiguity. I want the LLM to respond…

I'm using regular GPT-5, no custom instructions and memory turned off. Can you link to your shared Zealandia result? I think that mural result was spectacularly impressive, given that it started with a photo I took of the mural with almost no additional context.

I can't link since it's in an enterprise account.

Interestingly I tried the same question in a separate ChatGPT account and it gave a similar response you got. Maybe it was pulling context from the (separate) chat thread where it was talking about Zealandia. Which raises another question: once it gets something wrong once, will it just keep reenforcing the inaccuracy in future chats? That could lead to some very suboptimal behavior.

Getting back on topic, I strongly dislike the argument that this is all "user error". These models are on track to be worth a trillion dollars at some point in the future. Let's raise our expectations of them. Fix the models, not the users.

Re: An LLM is a lossy encyclopedia

#345
post #344
post #343

Earlier quoted context omitted.

I'm using regular GPT-5, no custom instructions and memory turned off. Can you link to your shared Zealandia result? I think that mural result was spectacularly impressive, given that it started with a photo I took of the mural with almost no additional context.

I can't link since it's in an enterprise account. Interestingly I tried the same question in a separate ChatGPT account and it gave a similar response you got. Maybe it was pulling context from the (separate) chat thread where it was talking about Zealandia. Which raises another question: once it gets something wrong once, will it just keep reenforcing the inaccuracy in future chats? That could lead to some very subo…

I wonder if you're stuck on an older model like GPT-4o?

EDIT: I think that's likely what is happening here: I tried the prompt against GPT-4o and got this https://chatgpt.com/share/68b8683b-09b0-8006-8f66-a316bfebda...

My consistent position on this stuff is that it's actually way harder to use than most people (and the companies marketing it) let on.

I'm not sure if it's getting easier to use over time either. The models are getting "better" but that partly means their error cases are harder to reason about, especially as they become less common.

Re: An LLM is a lossy encyclopedia

#346
post #337

Earlier quoted context omitted.

I think this actually points at a different problem, a problem with LLM users, but only to the extent that it's a problem with people with respect to any questions they have to ask any source they consider an authority at all. No LLM, nor any other source on the Internet, nor any other source off the Internet, can give you reliable dosage guidelines for copper peptides because this is information that is not known to…

> a problem with LLM users I think the flaw here is placing blame on users rather than the service provider. HN is cutting LLM companies slack because we understand the technical limitations making it hard for the LLM to just say “I don’t know”. In any other universe, we would be blaming the service rather than the user. Why don’t we fix LLMs so they don’t spit out garbage when it doesn’t know the answer. Have we giv…

> In any other universe, we would be blaming the service rather than the user.

I think the key question is "How is this service being advertised?"

Perhaps the HN crowd gives it a lot of slack because they ignore the advertising. Or if you're like me, aren't even aware of how this is being marketed. We know the limitations, and adapt appropriately.

I guess where we differ is on whether the tool is broken or not (hence your use of the word "fix"). For me, it's not at all broken. What may be broken is the messaging. I don't want them to modify the tool to say "I don't know", because I'm fairly sure if they do that, it will break a number of people's use cases. If they want to put a post-processor that filters stuff before it gets to the user, and give me an option to disable the post-processor, then I'm fine with it. But don't handicap the tool in the name of accuracy!

Re: An LLM is a lossy encyclopedia

#347
post #243

Earlier quoted context omitted.

Using a LLM for medical research is just as dangerous as Googling it. Always ask your doctors!

Not really: it's arguably quite a lot worse. Because you can judge the trustworthiness of the source when you follow a link from Google (e.g. I will place quite a lot of faith in pages at an .nhs.uk URL), but nobody knows exactly how that specific LLM response got generated.

Many of the big LLMs do RAG and will provide links to sources, eg. Bing/ChatGPT, Gemini Pro 2.5, etc.

Re: An LLM is a lossy encyclopedia

#348
post #337

Earlier quoted context omitted.

> a problem with LLM users I think the flaw here is placing blame on users rather than the service provider. HN is cutting LLM companies slack because we understand the technical limitations making it hard for the LLM to just say “I don’t know”. In any other universe, we would be blaming the service rather than the user. Why don’t we fix LLMs so they don’t spit out garbage when it doesn’t know the answer. Have we giv…

> In any other universe, we would be blaming the service rather than the user. I think the key question is "How is this service being advertised?" Perhaps the HN crowd gives it a lot of slack because they ignore the advertising. Or if you're like me, aren't even aware of how this is being marketed. We know the limitations, and adapt appropriately. I guess where we differ is on whether the tool is broken or not (hence…

The point you were making elsewhere in the thread was that "this is a bad use case for LLMs" ... "Don't use LLMs for dosing guidelines." ... "Using dosing guidelines is a bad example for demonstrating how reliable or unreliable LLMs are", etc etc etc.

You're blaming the user for having a bad experience as a result of not using the service "correctly".

I think the tool is absolutely broken, considering all of the people saying dosing guidelines is an "incorrect" use of LLM models. (While I agree it's not a good use, I strongly dislike how you're blaming the user for using it incorrectly - completely out of touch with reality).

We can't just cover up the shortfalls of LLMs by saying things like "Oh sorry, that's not a good use case, you're stupid if you use the tool for that purpose".

I really hope the HN crowd stops making excuses for why it's okay that LLMs don't perform well on tasks it's commonly asked to do.

> But don't handicap the tool in the name of accuracy!

If you're taking the position that it's the user's fault for asking LLMs a question it won't be good at answering, then you can't simultaneously advocate for not censoring the model. If it's the user's responsibility to know how to use ChatGPT "correctly", the tool (at a minimum) should help guide you away from using it in ways it's not intended for.

If LLMs were only used by smarter-than-average HN-crowd techies, I'd agree. But we're talking about a technology used by middle school kids. I don't think it's reasonable to expect middleschoolers to know what they should and shouldn't ask LLMs for help with.

Re: An LLM is a lossy encyclopedia

#349
post #194

Earlier quoted context omitted.

> The major component of what we call intelligence is purely predictive. Making more unfounded, nonsensical claims does not reinforce your first unfounded, nonsensical claim. I'm sure statisticians would love it if the human mind were an inference machine, but that doesn't make it one. Your point of view on this is faith-based.

His view aligns both with a leading neuroscience explanation of the brain (predictive coding [1]) and with Active Inference / the Free Energy principle [2] from optimal control theory. A similar theory of intelligence, called H-JEPA (hierarchical joint embedding predictive architecture) [3] is also put forward by Yann LeCun, a major AI pioneer. Another AI pioneer, Jürgen Schmidhuber, subsequently criticized LeCun's t…

The Bayesian brain model is an unfalsifiable, faith-based mechanism - as I alluded to in my previous comment.

Real science is done with it as a starting point, but it is not real science and claiming that it is an accurate representation of the human mind carries as much merit as claiming that "the soul" is what powers human intellect.

Re: An LLM is a lossy encyclopedia

#350

Earlier quoted context omitted.

That's not exactly true. Every time you start a new conversation; you get a new LLM for all intents. Asking an LLM about an unrelated topic towards the end of a ~500 page conversation will get you vastly different results than at the beginning. If we could get to multi-thousand page contexts, it would probably be less accurate than a human, tbh.

Yes, I should have clarified that I was referring to memory of training data, not of conversations.

Training data also deteriorates quite quickly as the context gets longer.
Post reply on HN