Live data from Hacker News

Local LLMs versus offline Wikipedia

evanhahn.com

101–110 of 200 posts

Re: Local LLMs versus offline Wikipedia

#101
post #51

Earlier quoted context omitted.

> LLMs will return faulty or imprecise information at times To be fair, so do humans and wikipedia.

It appears there's an expectation many non-tech people have that humans can be incorrect but refuse to hold LLMs to the same standard, despite warnings.

Because we still assume that computers are precise things that do what you tell them to do, and react in predictable(-ish) ways.

We don't know how to deal with a non-deterministic output from a computer.

Even here on HN you will see people whose world view is basically "LLMs are good and how dare you doubt them"

Re: Local LLMs versus offline Wikipedia

#102
post #33
post #9

Earlier quoted context omitted.

An unreliable computer treated as a god by a pre-information-age society sounds like a Star Trek episode.

Definitely sounds like a plausible and fun episode. On the other hand, real history if filled with all sorts of things being treated as a god that were much worse than "unreliable computer". For example, a lot of times it's just a human with malice. So how bad could it really get

Person god was not as scalable as AI god, so there's that.

Re: Local LLMs versus offline Wikipedia

#103
post #4

This is a sensible comparison. My "help reboot society with the help of my little USB stick" thing was a throwaway remark to the journalist at a random point in the interview, I didn't anticipate them using it in the article! https://www.technologyreview.com/2025/07/17/1120391/how-to-r... A bunch of people have pointed out that downloading Wikipedia itself onto a USB stick is sensible, and I agree with them. Wikipedi…

Someone should start a company selling USB sticks pre-loaded with lots of prepper knowledge of this type. In addition to making money, your USB sticks could make a real difference in the event of a global catastrophe. You could sell the USB stick in a little box which protects it from electromagnetic interference in the event of a solar flare or EMP. I suppose the most important knowledge to preserve is knowledge abo…

> Someone should start a company selling USB sticks pre-loaded with lots of prepper knowledge of this type.

It amuses me to no end that people think civilization will collapse but they will still have access to robotics and working computers to peruse USB sticks at their leisure.

Re: Local LLMs versus offline Wikipedia

#104
post #103

Earlier quoted context omitted.

Someone should start a company selling USB sticks pre-loaded with lots of prepper knowledge of this type. In addition to making money, your USB sticks could make a real difference in the event of a global catastrophe. You could sell the USB stick in a little box which protects it from electromagnetic interference in the event of a solar flare or EMP. I suppose the most important knowledge to preserve is knowledge abo…

> Someone should start a company selling USB sticks pre-loaded with lots of prepper knowledge of this type. It amuses me to no end that people think civilization will collapse but they will still have access to robotics and working computers to peruse USB sticks at their leisure.

It depends on your collapse threat model. In any case, my assumption is that serious preppers already have EMP-shielded laptops and solar panels for a SHTF scenario. And serious preppers are probably doing some datahoard as well. The point is that there are economies of scale in the datahoard. Most of the work of datahoard is identifying data worth hoarding, setting up your scripts, monitoring your webcrawler, etc. Once you've got a drive full of data, replicating that drive is comparatively easy. That's why it could make sense to start a business selling replicated drives.

Maybe there is room for an "all-in-one" product offering with an energy-efficient laptop, solar panel, and TBs of useful data, all protected in an EMP storage case for the event of solar flare.

Re: Local LLMs versus offline Wikipedia

#105
post #33
post #9

Earlier quoted context omitted.

An unreliable computer treated as a god by a pre-information-age society sounds like a Star Trek episode.

Definitely sounds like a plausible and fun episode. On the other hand, real history if filled with all sorts of things being treated as a god that were much worse than "unreliable computer". For example, a lot of times it's just a human with malice. So how bad could it really get

"Definitely sounds like a plausible and fun episode."

There were several original Star Trek episodes that explored this scenario. Not plausible. Actual.

"So how bad could it really get"

Watch Rodenberry's orginal Star Trek to get some ideas.

Re: Local LLMs versus offline Wikipedia

#106

I've found this amusing because right now i'm downloading `wikipedia_en_all_maxi_2024-01.zim` so i can use it with an LLM with pages extracted using `libzim` :-P. AFAICT the zim files have the pages as HTML and the file i'm downloading is ~100GB. (reason: trying to cross-reference my tons of downloaded games my HDD - for which i only have titles as i never bothered to do any further categorization over the years asid…

> trying to cross-reference my tons of downloaded games my HDD - for which i only have titles as i never bothered to do any further categorization over the years aside than the place i got them from - with wikipedia articles - assuming they have one - to organize them in genres, some info, etc and after some experimentation it turns out an LLM - specifically a quantized Mistral Small 3.2 - can make some sense of the chaos while being fast enough to run from scripts via a custom llama.cpp program

You can do this a lot easier with Wikidata queries, and that will also include known video games for which an English Wikipedia article doesn't exist yet.

Re: Local LLMs versus offline Wikipedia

#107
post #7

One important distinction is that the strength of LLMs isn't just in storing or retrieving knowledge like Wikipedia, it’s in comprehension. LLMs will return faulty or imprecise information at times, but what they can do is understand vague or poorly formed questions and help guide a user toward an answer. They can explain complex ideas in simpler terms, adapt responses based on the user's level of understanding, and…

   it’s in comprehension … what they can do is understand 
Well, no. The glaringly obvious recent example was the answer that Adolf Hitler could solve global warming.

My friend's car is perhaps the less polarizing example. It wouldn't start and even had a helpful error code. The AI answer was you need to replace an expensive module. Took me about five minutes with basic tools to come up with a proper diagnosis (not the expensive module). Off to the shop where they confirmed my diagnosis and completed the repair.

The car was returned with a severe drivability fault and a new error code. AI again helpfully suggested replace a sensor. I talked my friend through how to rule out the sensor and again AI was proven way off base in a matter of minutes. After I took it for a test drive I diagnosed a mechanical problem entirely unrelated to AI's answer. Off to the shop it went where the mechanical problem was confirmed, remedied, and the physically damaged part was returned to us.

AI doesn't comprehend anything. It merely regurgitates whatever information it's been able to hoover up. LLMs merely are glorified search engines.

Re: Local LLMs versus offline Wikipedia

#108

Earlier quoted context omitted.

Indeed. Ideally, you don't want to trust other people's summaries of sources, but you want to look at the sources yourself, often with a critical eye. This is one of the things that everyone gets taught in school, everyone's says they agree with, and then just about no one does (and at times, people will outright disparage the idea). Once out of school, tertiary sources get treated as if they're completely reliable.…

A trustless society can't progress/function a lot. I trust doctors who treat me, civil engineers who built my house and even in software which I pretend to be expert in I haven't seen source code of any OS and browser I use as I trust on companies or OSS devs. Most of this is based on reputation. LLMs are same, I just have to calculate level of trust as I use it.

Some trust is necessary, yes, but not complete trust. I certainly don't trust my coworkers code. I don't trust their services to return what they say they will return 100% of the time. I don't trust that someone won't introduce a bug.

I assert assumptions and dive into their code when something is fishy.

I also know nothing about health, but I'm going to double check what my doctors say. Maybe against a 2nd doctor, maybe against the Internet, or maybe just listen to what my body is saying. Doctors are frequently wrong. It's kind of astonishing and scary how much they don't know

Tldr trust but verify.

Re: Local LLMs versus offline Wikipedia

#109
post #51

Earlier quoted context omitted.

It appears there's an expectation many non-tech people have that humans can be incorrect but refuse to hold LLMs to the same standard, despite warnings.

On average, it is reasonable to expect that wikipedia will be more correct than an LLM

Doubtful.

Re: Local LLMs versus offline Wikipedia

#110
post #9
post #7

One important distinction is that the strength of LLMs isn't just in storing or retrieving knowledge like Wikipedia, it’s in comprehension. LLMs will return faulty or imprecise information at times, but what they can do is understand vague or poorly formed questions and help guide a user toward an answer. They can explain complex ideas in simpler terms, adapt responses based on the user's level of understanding, and…

An unreliable computer treated as a god by a pre-information-age society sounds like a Star Trek episode.

Surprised nobody has pointed out that this was and episode of the Twilight Zone [0], if you substitute "pre-information-age" with "post-information-age".

0. https://en.wikipedia.org/wiki/The_Old_Man_in_the_Cave

Post reply on HN