Live data from Hacker News

Pirate Library Mirror: Preserving 7TB of books (that are not in Libgen)

pilimi.org

171–180 of 438 posts

Re: Pirate Library Mirror: Preserving 7TB of books (that are not in Libgen)

#171

Earlier quoted context omitted.

OK, what's the napkin math?

I don't know about "all the knowledge ever", but to give you a baseline, the entire Wikipedia with images but without editing history, as last archived by Kiwix in May, 2022, is 90 Gb. A web server can run just fine on e.g. Raspberry Pi Zero W ($10), exposing any such content to any smartphone etc able to connect to it via WiFi (Kiwix sells preconfigured SD cards for their content, even). So, assuming that most peopl…

There are some language models like DeepMind RETRO that can make use of a 1TB text collection after the model is trained. The idea is to chunk up the text and make the blocks searchable with a neural embedding index. When you ask a question to the model, it first searches for relevant information and then adds it to the prompt. The result - you can get GPT-3 quality on a 25x smaller model. That means you could have your own GPT-3. Another advantage is that you can reference the sources and the generated text is more grounded.

An open source project could make a "search-engine-language-model" and in turn make this library much more accessible.

A recent article related to this: http://mitchgordon.me/ml/2022/07/01/retro-is-blazing.html

Re: Pirate Library Mirror: Preserving 7TB of books (that are not in Libgen)

#172

Earlier quoted context omitted.

No one would make money off of it. That is the point. There are innumerable things which are worthwhile which are not profitable. The idea that we do not do a world library of free digital copies of every book ever written really highlights the problem with the thinking your comment has demonstrated. The idea that individual pursuits must make a profit to be justified leads to us doing terrible things like: not makin…

> libertarian society as long as we have community ownership of the means of production Libertarianism includes the right to own property. Also common ownership of the means of production is an old idea, and has been tried many times. It always results in poverty.

Ah hello Walter. I recognize your username as you and I have disagreed on this before.

I’m not saying people shouldn’t have the right to own property. But that a good way of organizing society is collective ownership of the means of production. If you are part owner in something with shares and a contract, that obviously still relies on property rights. That is how the stock market works after all.

EDIT: I am basically proposing a change in norms, rather than a change in rights. Currently the norm is individual private owners or ownership by a board of directors. I am proposing ownership by communities as a collective. Same rights involved, but a different norm.

Re: Pirate Library Mirror: Preserving 7TB of books (that are not in Libgen)

#173
post #163

Earlier quoted context omitted.

No one would make money off of it. That is the point. There are innumerable things which are worthwhile which are not profitable. The idea that we do not do a world library of free digital copies of every book ever written really highlights the problem with the thinking your comment has demonstrated. The idea that individual pursuits must make a profit to be justified leads to us doing terrible things like: not makin…

You're responding to a comment which aims to make your same point.

Ah. Reading it again, I can see what you mean. Unfortunately I encounter a lot of people here who would write that comment in all seriousness, so if it was sarcastic I hadn’t noticed. I think I’m still not sure about that.

Re: Pirate Library Mirror: Preserving 7TB of books (that are not in Libgen)

#174

Earlier quoted context omitted.

Wasn't that what google books was supposed to be?

Google Books, like so many Google projects, had a dual purpose. Making books accessible is noble and on-mission. But more importantly natural language models can be trained on the scanned corpus. The same was true of the original GOOG 411, which provided a free service, but was really put in place to train up their voice recognition projects. This is a long running strategy of Google, and it's a shrewd one. The main…

That can't be true, Google Books was 15 years prior to the advent of large language models. Until 2020 nobody could train on such a large collection.

I think Google initially wanted to augment the web results with a large book collection to get "all the world information and make it searchable", same with Google News.

Re: Pirate Library Mirror: Preserving 7TB of books (that are not in Libgen)

#175
post #124

Again, it's just insane to me that we don't even much have a meaningful discussion of: "Hey, wait, literally everyone could have the entire library of Alexandria in their house for a couple hundred bucks per person. Like, all the knowledge ever. Maybe that should be considered the good default of things. At least one in every town that everyone could use, for free, forever, without restriction to ANY of the knowledge…

"The African economy will start booming once they get UNLIMITED access to those 1800s manuals on measuring the Ether and where the best seal clubbing sites are in Alaska." Nobody ever thinks this through. Books are basically worthless at this point

then maybe we stop letting people push the copyright periods tentatively so more relevant things come into the public domain as well. this is a nothing argument.

Re: Pirate Library Mirror: Preserving 7TB of books (that are not in Libgen)

#176
post #124

Again, it's just insane to me that we don't even much have a meaningful discussion of: "Hey, wait, literally everyone could have the entire library of Alexandria in their house for a couple hundred bucks per person. Like, all the knowledge ever. Maybe that should be considered the good default of things. At least one in every town that everyone could use, for free, forever, without restriction to ANY of the knowledge…

Insane in the abstract, but not exactly unfathomable. Who would make money off of that, and who else makes money off of it's absense?

Imagine for a split second an idea of value that doesn't start and end with private profit

Re: Pirate Library Mirror: Preserving 7TB of books (that are not in Libgen)

#177
post #124

Again, it's just insane to me that we don't even much have a meaningful discussion of: "Hey, wait, literally everyone could have the entire library of Alexandria in their house for a couple hundred bucks per person. Like, all the knowledge ever. Maybe that should be considered the good default of things. At least one in every town that everyone could use, for free, forever, without restriction to ANY of the knowledge…

That's already the case minus the last ~70 years or so. The overwhelming majority of our knowledge is in the public domain, in particular cultural artifacts.

It's a nice sentiment but like, people can already go to gutenberg.org and download pretty much most important works of literature in existence and most books have like 5k downloads so there's that.

Re: Pirate Library Mirror: Preserving 7TB of books (that are not in Libgen)

#178

Earlier quoted context omitted.

> libertarian society as long as we have community ownership of the means of production Libertarianism includes the right to own property. Also common ownership of the means of production is an old idea, and has been tried many times. It always results in poverty.

Ah hello Walter. I recognize your username as you and I have disagreed on this before. I’m not saying people shouldn’t have the right to own property. But that a good way of organizing society is collective ownership of the means of production. If you are part owner in something with shares and a contract, that obviously still relies on property rights. That is how the stock market works after all. EDIT: I am basical…

The US is a free country. You can form a voluntary collective any time you like.

Re: Pirate Library Mirror: Preserving 7TB of books (that are not in Libgen)

#179
post #142

Earlier quoted context omitted.

Society as a whole. Exactly the kind of initiatives governments are supposed to be taking. Instead our governments are totally captured by the profit-seeking organizations and are hellbent on using their monopoly of violence to imprison activists working on these initiatives

Interesting difference here between how YC reacts to Codex vs this library. Both are reusing copyrighted materials to help society.

maybe it has something to do with the commercial nature of one of those two things.

Re: Pirate Library Mirror: Preserving 7TB of books (that are not in Libgen)

#180
post #15

It's really funny to think about how the advances of technology keeps changing how we perceive books. 7TB is even a commodity disk these days. And it's a lot less than the torrent of scientific papers that floated around some time ago (that was ~18TB IIRC).

It's 7TB compressed. If it's text you'd need about 70TB to decompress it. It's probably mostly images though, so probably not quite that bad.

Things can be rendered from compressed container files. For HTML with images, even slow-but-strong compression like LZMA is already fast enough to render pages as fast as you can click through them, even on fairly old hardware.

Kiwix .ZIM file format is a good example. The entire Gutenberg Library is a single ~65 Gb file, and you can read any book from it without unpacking anything.

Post reply on HN