Live data from Hacker News

Having a 20GB file that lets you ask an offline computer any question is amazing

old.reddit.com

41–50 of 127 posts

Re: Having a 20GB file that lets you ask an offline computer any question is amazing

#41

We can do that, but then we will need to close all applications to run it because it will need 20GB RAM or VRAM to run. It is still more practical to visit stackoverflow when I have a question. (still impressive how much this improves every week. Maybe in 2-3 months we have something more useful)

I just upgraded my 2020 era thinkpad (t470, which I bought this time last year for less than £300) to 32g memory for £70. 20g of ram doesn’t feel like a lot today.

Also today: soldered memory and 8GB MBPs that START at $1300.

Re: Having a 20GB file that lets you ask an offline computer any question is amazing

#42
post #25

Earlier quoted context omitted.

Given the biased nature of many if not most Wikipedia articles it would be certain that such a system would bend the truth to an unacceptable degree. Train the thing on more than just Wikipedia, e.g add Encyclopedia Britannica to the training set. This is the one area where "diversity" really matters: diversity of opinion. Using a single biased source as your oracle will turn you into a pawn for those who control the…

what? since when wikipedia is biased? and about what? (and why EB isn't?)

Example: https://nitter.net/echetus/status/1579776106034757633#m

Re: Having a 20GB file that lets you ask an offline computer any question is amazing

#43
post #21

I currently have a copy of offline Wikipedia (and Stack Overflow, WikiVoyage, etc.) via Kiwix on my laptop and also LLaMA-style models at the same time (since I like things being local), and with respect to asking an LLM questions, it would be great if it could leverage the offline Wikipedias to ensure it stays grounded and not make up facts.

How much space does that take up, especially SO? There's so much junk on there.

Re: Having a 20GB file that lets you ask an offline computer any question is amazing

#44
post #25

Earlier quoted context omitted.

what? since when wikipedia is biased? and about what? (and why EB isn't?)

US politics, definitely. But also any other topic that attracts an internet mob large enough to infiltrate the editor community. Take the first few paragraphs of Donald Trump [0] vs. Joe Biden [1] and try to compare the way they're presented: Donald John Trump [...] and his businesses have been involved in more than 4,000 state and federal legal actions, including six business bankruptcies. Trump promoted conspiracy…

The article covers allegations of corruption and mentions his "gaffes" a number of times.

Sometimes where there's no smoke, there's no fire. The guy is very boring and has a weird kid. That's about it.

Re: Having a 20GB file that lets you ask an offline computer any question is amazing

#45
post #25

Earlier quoted context omitted.

Given the biased nature of many if not most Wikipedia articles it would be certain that such a system would bend the truth to an unacceptable degree. Train the thing on more than just Wikipedia, e.g add Encyclopedia Britannica to the training set. This is the one area where "diversity" really matters: diversity of opinion. Using a single biased source as your oracle will turn you into a pawn for those who control the…

what? since when wikipedia is biased? and about what? (and why EB isn't?)

Wikipedia is as biased as the writers of the articles. Wikipedia has articles outlining bias in its own articles. These things are authored by people, and people are biased. Wikipedia knows this and so should its readers. They aren't unique.

Re: Having a 20GB file that lets you ask an offline computer any question is amazing

#46
post #39

Earlier quoted context omitted.

I would pay several hundred dollars a month to keep access to GPT-4, and never found any use at all for crypto. There might be hype, but there’s also an insane amount of utility

Some examples please. I’ve yet to find any actual utility, mostly I’ve got a lot of convincing bullshit out of it.

I use it to speed up writing boilerplate code. That alone saves me enough time to justify paying for access many times over.

Re: Having a 20GB file that lets you ask an offline computer any question is amazing

#47
post #25

Earlier quoted context omitted.

Given the biased nature of many if not most Wikipedia articles it would be certain that such a system would bend the truth to an unacceptable degree. Train the thing on more than just Wikipedia, e.g add Encyclopedia Britannica to the training set. This is the one area where "diversity" really matters: diversity of opinion. Using a single biased source as your oracle will turn you into a pawn for those who control the…

what? since when wikipedia is biased? and about what? (and why EB isn't?)

[deleted]

Re: Having a 20GB file that lets you ask an offline computer any question is amazing

#48
post #25

Earlier quoted context omitted.

what? since when wikipedia is biased? and about what? (and why EB isn't?)

US politics, definitely. But also any other topic that attracts an internet mob large enough to infiltrate the editor community. Take the first few paragraphs of Donald Trump [0] vs. Joe Biden [1] and try to compare the way they're presented: Donald John Trump [...] and his businesses have been involved in more than 4,000 state and federal legal actions, including six business bankruptcies. Trump promoted conspiracy…

Exactly. This is the tip of the iceberg.

Re: Having a 20GB file that lets you ask an offline computer any question is amazing

#49
post #43
post #21

I currently have a copy of offline Wikipedia (and Stack Overflow, WikiVoyage, etc.) via Kiwix on my laptop and also LLaMA-style models at the same time (since I like things being local), and with respect to asking an LLM questions, it would be great if it could leverage the offline Wikipedias to ensure it stays grounded and not make up facts.

How much space does that take up, especially SO? There's so much junk on there.

Around 159 GB, it's surprisingly small.

Re: Having a 20GB file that lets you ask an offline computer any question is amazing

#50

Earlier quoted context omitted.

I also don’t want my AI in the cloud favoring a corporation’s values and goals. But these tools are almost certainly going to come from them. And they want you to store all your data in the cloud too. And they want to keep you from running the software you want.

I think there is one company very well placed to do this, Apple. Their revenue is driven by hardware sales, rather than advertising, and they have committed to some level of "privacy" focussed design (although obviously there are criticisms). On top of that their hardware is looking almost like it was design for exactly these use cases. Apple silicone with its unified memory architecture, GPU and Neural cores is very…

> Apple silicone with its unified memory architecture, GPU and Neural cores is very well placed for local LLMs.

Yes, university research from 2011 already showed a privacy focused design to machine learning[1]. But will this be private or public like Wikipedia, Bitcoin and Bittorrent?

> Effectively an extension to your own brain.

Correct. After 1996 fantasy/visionary publications about this, we now know roughly how to do this. Most difficult problem is who should own it, no corporation in control is desired. That means full decentralisation, federated learning is not enough.

But fully decentralised learning is hard. Try to apply machine learning in permissionless, byzantine, unsupervised, decentralised, adversarial, continuous learning context. See my lab at Delft University focused for a decade already on "The Global Brain"[2]. With 2.3 million download and crowd-sourcing, we might get there..

[1] "Gossip Learning with Linear Models on Fully Distributed Data", 2011, https://arxiv.org/abs/1109.1396

[2] The Global Brain - the roadmap, https://github.com/Tribler/tribler/issues/7064

Post reply on HN