Earlier quoted context omitted.
I gave LLM a list of python packages and asked it to give me their respective licenses. Obviously it got some of them wrong. I had to manually check with the package's pypi page.
Ok? Yes. That is the problem; that sometimes it works. See the topic. Adding RAG or web search capability limits the loss and hallucinations. Yes. You always need to check the results. Your task by the way is better for an agentic AI system that can web search, get, and double check results.
An LLM is a lossy encyclopedia
331–340 of 365 posts
Re: An LLM is a lossy encyclopedia
#332Earlier quoted context omitted.
Yeah an LLM is an unreliable librarian, if anything.
That’s a much better analogy. You have to specifically ask them for information and they will happily retrieve it for you, but because they are unreliable they may get you the wrong thing. If you push back they’ll apologise and try again (librarians try to be helpful) but might again give you the wrong thing (you never know, because they are unreliable).
A librarian might bring you the wrong book, that's the former. An LLM does the latter. They are not the same.
Re: An LLM is a lossy encyclopedia
#333Earlier quoted context omitted.
No, I don't think GPT-5 clarifying questions actually do what you think they do. They just made the model ask clarifying questions for the sake of asking clarifying questions. I'm sure GPT-4o would have given me the answer I wanted without clarifying questions.
revisit your instructions.md and/or user preferences, this is very likely the root cause
Re: An LLM is a lossy encyclopedia
#334Earlier quoted context omitted.
Sorry, 2 billion years of neurobiology beats 60 years of NLP/LLMs which knows less to nothing about language since "arbitrary points can never be refined or defined to specifics" check your corners and know your inputs. The bill is due on NLP.
Incoherent drivel.
Re: An LLM is a lossy encyclopedia
#335Earlier quoted context omitted.
> I used ChatGPT 5 over the weekend to double check dosing guidelines for a specific medication. This use case is bad by several degrees . Consider an alternative: Using Google to search for it and relying on its AI generated answer. This usage would be bad by one degree less, but still bad. What about using Google and clicking on one of the top results? Maybe healthline.com? This usage would reduce the badness by on…
The compound I was researching was [edit: removed]. Problem is it's not FDA approved, only prescribed by compounding pharmacies off label. Experimental compound with no official guidelines. The first result on Google for "[edit: removed] dosing guidelines" is a random word document hosted by a Telehealth clinic. Not exactly the most reliable source. Edit: Jeesh, what’s with the downvotes?
It's an uncomfortable position to be in trying to biohack your way to a more youthful appearance using treatments that have never been studied in human trials, but that's the reality you're facing. Whatever guidelines you manage to find, whether from the telehealth clinic directly, or from a language model that read the Internet and ingested that along with maybe a few other sources, are generally extrapolated from early rodent studies and all that's being extrapolated is an allometric scaling from rat body to human body of the dosage the researchers actually gave to the rats. What effect that actually had, and how that may or may not translate to humans, is not usually a part of the consideration. To at least some extent, it can't be if the compound was never trialed on humans.
You're basically just going with scale up a dosage to human sized that at least didn't kill the rats. Take that and it probably won't kill you. What it might actually do can't be answered, not by doctors, not by an LLM, not by Wikipedia, not by anecdotes from past biohackers who tried it on themselves. This is not a failure of information retrieval or compression. You're just asking for information that is not known to anyone, so no one can give it to you.
If there's a problem here specific to LLMs, it's that they'll generally give you an answer anyway and will not in any way quantify the extent to which it is probably bullshit and why.
Re: An LLM is a lossy encyclopedia
#336Earlier quoted context omitted.
That’s a much better analogy. You have to specifically ask them for information and they will happily retrieve it for you, but because they are unreliable they may get you the wrong thing. If you push back they’ll apologise and try again (librarians try to be helpful) but might again give you the wrong thing (you never know, because they are unreliable).
There's a big difference between giving you correct information about the wrong thing, vs giving you incorrect information about the right thing. A librarian might bring you the wrong book, that's the former. An LLM does the latter. They are not the same.
Re: An LLM is a lossy encyclopedia
#337Earlier quoted context omitted.
The compound I was researching was [edit: removed]. Problem is it's not FDA approved, only prescribed by compounding pharmacies off label. Experimental compound with no official guidelines. The first result on Google for "[edit: removed] dosing guidelines" is a random word document hosted by a Telehealth clinic. Not exactly the most reliable source. Edit: Jeesh, what’s with the downvotes?
I think this actually points at a different problem, a problem with LLM users, but only to the extent that it's a problem with people with respect to any questions they have to ask any source they consider an authority at all. No LLM, nor any other source on the Internet, nor any other source off the Internet, can give you reliable dosage guidelines for copper peptides because this is information that is not known to…
I think the flaw here is placing blame on users rather than the service provider.
HN is cutting LLM companies slack because we understand the technical limitations making it hard for the LLM to just say “I don’t know”.
In any other universe, we would be blaming the service rather than the user.
Why don’t we fix LLMs so they don’t spit out garbage when it doesn’t know the answer. Have we given up on that thought?
Re: An LLM is a lossy encyclopedia
#338Earlier quoted context omitted.
I think this actually points at a different problem, a problem with LLM users, but only to the extent that it's a problem with people with respect to any questions they have to ask any source they consider an authority at all. No LLM, nor any other source on the Internet, nor any other source off the Internet, can give you reliable dosage guidelines for copper peptides because this is information that is not known to…
> a problem with LLM users I think the flaw here is placing blame on users rather than the service provider. HN is cutting LLM companies slack because we understand the technical limitations making it hard for the LLM to just say “I don’t know”. In any other universe, we would be blaming the service rather than the user. Why don’t we fix LLMs so they don’t spit out garbage when it doesn’t know the answer. Have we giv…
Re: An LLM is a lossy encyclopedia
#339Earlier quoted context omitted.
> a problem with LLM users I think the flaw here is placing blame on users rather than the service provider. HN is cutting LLM companies slack because we understand the technical limitations making it hard for the LLM to just say “I don’t know”. In any other universe, we would be blaming the service rather than the user. Why don’t we fix LLMs so they don’t spit out garbage when it doesn’t know the answer. Have we giv…
Current frontier LLMs - Claude 4, GPT-5, Gemini 2.5 - are massively more likely to say "I don't know" than last year's models.
What’s an example prompt where it will say “idk”?
Edit: Just tried a silly one, asking it to tell me about the 8th continent on earth, which doesn’t exist. How difficult is it for the model to just say “sorry, there are only 7 continents”. I think we should expect more from LLMs and stop blaming things on technical limitations. “It’s hard” is getting to be an old excuse considering the amount of money flowing into building these systems.
Re: An LLM is a lossy encyclopedia
#340I totally agree with the author. Sadly, I feel like that's not what the majority of LLM users tend to view LLMs. And it's definitely not what AI companies marketing. > The key thing is to develop an intuition for questions it can usefully answer vs questions that are at a level of detail where the lossiness matters the problem is that in order to develop an intuition for questions that LLMs can answer, the user will…
I can't think of any other tools like this. An LLM can multiply your efforts, but only if you were capable of doing it yourself. Wild.