An LLM is a lossy encyclopedia
simonwillison.net
An LLM is a lossy encyclopedia
1–10 of 365 posts
Re: An LLM is a lossy encyclopedia
#2Re: An LLM is a lossy encyclopedia
#3Imagine a slightly lossy compression algorithm which can store 10x, 100x the current best lossless and be able to maintain 99.999% fidelity when recalling that information. Probably, very probably a pipe dream. But why do large on device models seem to be able to remember adjust everything from Wikipedia and store that in smaller format than a direct archive of the source Material. (Look at the current best from diffusion models as well)
Re: An LLM is a lossy encyclopedia
#4llm is a pretty good librarian who has read a ton of books (and doesn't have perfect memory)
even more useful when allowed to think-aloud
even more useful when allowed to write stuff down and check in library db
even more useful when allowed to go browse and pick up some books
even more useful when given a budget for travel and access to other archives
even more useful when …
brrrrt
Re: An LLM is a lossy encyclopedia
#5The models hold more information than they can immediately extract, but CoT can find a key to look it up or synthesise by applying some learned generalisations.
Re: An LLM is a lossy encyclopedia
#6I've got my opinion on whether that's useful or not and it's quite a bit more nuanced. You don't zoom-enhance JPEGs for a reason either.
Re: An LLM is a lossy encyclopedia
#7Re: An LLM is a lossy encyclopedia
#8Re: An LLM is a lossy encyclopedia
#9Re: An LLM is a lossy encyclopedia
#10I disagree with that analogy, because LLMs have a lot of connections between text fragments, which an encyclopedia doesn't have to such a deep degree. An encyclopedia also can't interpret and output relevant knowledge from an input prompt.
A slightly more precise analogy is probably 'a lossily compressed snapshot of the web'. Or maybe the Librarian from Snow Crash - but at least that one knew when it didn't know ;)