Live data from Hacker News

Local LLMs versus offline Wikipedia

evanhahn.com

131–140 of 200 posts

Re: Local LLMs versus offline Wikipedia

#131
post #75
post #9

Earlier quoted context omitted.

An unreliable computer treated as a god by a pre-information-age society sounds like a Star Trek episode.

Remember the first time you touched a computer, the first game you ever played or the first little script you wrote that did something useful. I imagine this is how a lot of people feel when using LLM's especially now that it's new. It is the most incredible technology ever created by this point in our history imo and the cynicism on HN is astounding to me.

Maybe it’s just psychology at work, but I see a huge difference between that time 15 years ago when I wrote my first useful script, and that time last week when an LLM spat out a piece of code to solve an issue I had.

The former made me so proud. My learning had paid off, and maybe there was nothing I couldn’t do. I had laid my pattern of thought onto the machine and made it do my bidding through sheer logic and focus. I had unlocked something special.

The latter was just OpenAI opaquely doing stuff for me while I watched a TV show in the background. No focus or logic was really necessary. I probably learned something from this, but not nearly as much as I could’ve if I actually read the docs and tried it myself.

I’ve also dabbled in art and design over the years, and I recognise this as the same difference as between painting something you’re truly proud of and asking Midjourney to generate you some images.

Then again, maybe that’s just how technological progress works. My great-great-grandmother was probably really proud and happy when she sewed and embroidered a beautiful shirt, but my shirts come from a store and I don’t really think about it.

Re: Local LLMs versus offline Wikipedia

#132
post #75
post #9

Earlier quoted context omitted.

An unreliable computer treated as a god by a pre-information-age society sounds like a Star Trek episode.

Remember the first time you touched a computer, the first game you ever played or the first little script you wrote that did something useful. I imagine this is how a lot of people feel when using LLM's especially now that it's new. It is the most incredible technology ever created by this point in our history imo and the cynicism on HN is astounding to me.

I have been involved with AI for over 40 years. I assure you anyone being shown a current frontier model in operation 10 years ago would have been blown off their socks.

Yet here we are. Rather than exploring this fantastic new tool, so many here are obsessed with pointing out flaws and shortcomings.

I get the angst of a world facing dramatic change. I don't get the denial and deliberate ignorance flaunted as somehow deep insight.

Re: Local LLMs versus offline Wikipedia

#133
post #4

This is a sensible comparison. My "help reboot society with the help of my little USB stick" thing was a throwaway remark to the journalist at a random point in the interview, I didn't anticipate them using it in the article! https://www.technologyreview.com/2025/07/17/1120391/how-to-r... A bunch of people have pointed out that downloading Wikipedia itself onto a USB stick is sensible, and I agree with them. Wikipedi…

I've been carrying around a local wikipedia dump on my phone or pda for quite a bit more than 10 years now (including with pictures for the last 5 years). Before kiwix and zim, I used tomeraider and aard.

I do it both for disaster preparedness but also off-line preparedness. Happens more often than you'd think.

But I have been thinking about how useful some of the models are these days, and the obvious next step to me seems to be to pair a local model with a local wikipedia in a RAG style set up so you get the best of both.

Re: Local LLMs versus offline Wikipedia

#134
post #51

Earlier quoted context omitted.

> LLMs will return faulty or imprecise information at times To be fair, so do humans and wikipedia.

It appears there's an expectation many non-tech people have that humans can be incorrect but refuse to hold LLMs to the same standard, despite warnings.

We have decades of experience with computers being deterministic machines that will return a correct output given a correct input and program.

I can’t multiply large numbers in my head, but if I plug 273*8113 into a calculator, I can expect it to give me the same, correct answer every time.

Now suddenly it’s „Well yes, it can make mistakes, but so can humans! Sometimes it’ll be right, but also sometimes it’ll make up a random answer, kinda like humans!”, which I suppose is true, but it’s also nonsense - the very reason I was using technology (in that case, a calculator) to do my work is because I wanted to avoid mistakes that a human (me) would make without it. If a piece of tech can’t be reliably expected to perform a task better than a person can on their own, then what’s really the point?

Re: Local LLMs versus offline Wikipedia

#135
post #75

Earlier quoted context omitted.

Remember the first time you touched a computer, the first game you ever played or the first little script you wrote that did something useful. I imagine this is how a lot of people feel when using LLM's especially now that it's new. It is the most incredible technology ever created by this point in our history imo and the cynicism on HN is astounding to me.

I have been involved with AI for over 40 years. I assure you anyone being shown a current frontier model in operation 10 years ago would have been blown off their socks. Yet here we are. Rather than exploring this fantastic new tool, so many here are obsessed with pointing out flaws and shortcomings. I get the angst of a world facing dramatic change. I don't get the denial and deliberate ignorance flaunted as somehow…

Sure. There is also still a massive chasm between those frontier models and what the hype is pushing too.

Re: Local LLMs versus offline Wikipedia

#136
post #7

One important distinction is that the strength of LLMs isn't just in storing or retrieving knowledge like Wikipedia, it’s in comprehension. LLMs will return faulty or imprecise information at times, but what they can do is understand vague or poorly formed questions and help guide a user toward an answer. They can explain complex ideas in simpler terms, adapt responses based on the user's level of understanding, and…

Understanding the question is more valuable than giving the correct answer?

That’s the basis of a cult.

Re: Local LLMs versus offline Wikipedia

#137

Earlier quoted context omitted.

On average, it is reasonable to expect that wikipedia will be more correct than an LLM

Doubtful.

How? The LLM is trained on the same information. In a lossy way, I might add. So how on Earth would the LLM be as reliable, let alone even more so?

Re: Local LLMs versus offline Wikipedia

#138

Earlier quoted context omitted.

Of course that’s angle they decide to open the article from. That they feel the need to frame these tools using the most grandiose terms bothers me. How does it make you feel?

I was once interviewed by my country's biggest paper about "strava art" I make, aka biking/running with a gps logger in order to create some kind of figure on the map. It was edited into this video about people drawing dicks on maps using this technique. Aka the intro was loads of penises on maps, and then "someone that enjoys making this kind of art is Mats here" and then the video interview started. When they ask w…

I’ve had a very similar experience. I was only on TV once. Right before Christmas, ~20 years ago, I was running some errands downtown and ran into a camera crew doing a puff piece about holiday preparations.

They asked me what was most important to me about the holidays, and I said that I really don’t care about the presents, but I love the atmosphere, the music, and spending time with my loved ones.

A couple days later the segment was aired, and it went something like this:

>Reporter: “Our crew asked people on the street what they like most about the holidays.”

>Teenage me: “…the presents…”

Re: Local LLMs versus offline Wikipedia

#139
post #9
post #7

One important distinction is that the strength of LLMs isn't just in storing or retrieving knowledge like Wikipedia, it’s in comprehension. LLMs will return faulty or imprecise information at times, but what they can do is understand vague or poorly formed questions and help guide a user toward an answer. They can explain complex ideas in simpler terms, adapt responses based on the user's level of understanding, and…

An unreliable computer treated as a god by a pre-information-age society sounds like a Star Trek episode.

Eh good enough, a better alternative when the elder/leader can't help than the alternative of asking the Pythia at Delphi

Re: Local LLMs versus offline Wikipedia

#140

I've found this amusing because right now i'm downloading `wikipedia_en_all_maxi_2024-01.zim` so i can use it with an LLM with pages extracted using `libzim` :-P. AFAICT the zim files have the pages as HTML and the file i'm downloading is ~100GB. (reason: trying to cross-reference my tons of downloaded games my HDD - for which i only have titles as i never bothered to do any further categorization over the years asid…

> trying to cross-reference my tons of downloaded games my HDD - for which i only have titles as i never bothered to do any further categorization over the years aside than the place i got them from - with wikipedia articles - assuming they have one - to organize them in genres, some info, etc and after some experimentation it turns out an LLM - specifically a quantized Mistral Small 3.2 - can make some sense of the…

I'm not sure about this, i just checked Tron 2.0 (just a random game i thought of) and Wikidata seems to have wrong info (e.g. genre) compared to the Wikipedia article. Also i need to it describe a bit with what the game is about since i want to generate an html file with all the games and do a quick scan of them and Wikidata doesn't have that.

IGDB would be a better source than Wikidata (especially since it does have a small description too) but i wanted to do things offline. And having Wikipedia locally doesn't hurt. And TBH i don't think it'd be any easier, extracting the data from Wikipedia pages was the most trivial part.

That said I'll need to use some other source at some point since, as you mentioned, Wikipedia does not have everything.

Post reply on HN