Live data from Hacker News

Pirate Library Mirror: Preserving 7TB of books (that are not in Libgen)

pilimi.org

211–220 of 438 posts

Re: Pirate Library Mirror: Preserving 7TB of books (that are not in Libgen)

#211

Earlier quoted context omitted.

I don't know about "all the knowledge ever", but to give you a baseline, the entire Wikipedia with images but without editing history, as last archived by Kiwix in May, 2022, is 90 Gb. A web server can run just fine on e.g. Raspberry Pi Zero W ($10), exposing any such content to any smartphone etc able to connect to it via WiFi (Kiwix sells preconfigured SD cards for their content, even). So, assuming that most peopl…

There are some language models like DeepMind RETRO that can make use of a 1TB text collection after the model is trained. The idea is to chunk up the text and make the blocks searchable with a neural embedding index. When you ask a question to the model, it first searches for relevant information and then adds it to the prompt. The result - you can get GPT-3 quality on a 25x smaller model. That means you could have y…

The interesting aspect of GPT-3+ is their reasoning like behavior such as Chains of thought or "step by step" generation. There's no evidence RETRO as is has sufficient capacity for the actually interesting behaviors.

Re: Pirate Library Mirror: Preserving 7TB of books (that are not in Libgen)

#212
post #209
post #185

Earlier quoted context omitted.

Hasn't been more knowledge published in the last 70 years than during all the times before? More than 2 million new books get published every year.

I assume they are referring to the fact that books older than about 70 years are in the public domain. The rest are protected by copyright and not free.

I think so too, but I am referring to this statement:

> The overwhelming majority of our knowledge is in the public domain

Re: Pirate Library Mirror: Preserving 7TB of books (that are not in Libgen)

#213
post #210

Excellent. I hope more and more people start to see how absurd and evil the concept of "intellectual property" is. It should be totally rejected and any form of keeping useful information for yourself should be shunned and tabooized. In todays world, many have been programmed to believe that the would could not exits without such immoral restrictions, which is horrible.

There should be a fine line in intellectual property rights. I see where you are coming from - quite often intellectual property is used a a moat to protect insane revenues and, as a repercussion, delay or slowdown our progress as humanity.

But it is also use to protect unique creator revenue and encourage to create more.

If you ask where the fine line should be I have no immediate answer, but abolishing intellectual property rights just like enforcing them at all costs doesn't seem to be the optimal course of action to me.

Re: Pirate Library Mirror: Preserving 7TB of books (that are not in Libgen)

#214

Earlier quoted context omitted.

(How) would you filter junk?

You'd trust the publisher, so you'd say "I want to help the Internet Archive with its archiving" and they'd be responsible for what files they pushed to your storage. This is basically an opportunistic, distributed filesystem that's designed to work with high latency links. Only the tracker can write to its nodes, but anyone can read.

This, and easy-to-configure bandwidth throttling was sorely missing from IPFS last time I tried it.

Being able to specify "use 50% of storage to cache low-traffic things, and the other 50% for high-traffic things" would be amazing.

Re: Pirate Library Mirror: Preserving 7TB of books (that are not in Libgen)

#215

Earlier quoted context omitted.

I found a book I've been looking for, but it was only an image scan pdf. I OCR'ed it, and I'm slowly fixing the errors and converting to epub. Tedious, but interesting.

Kudos to you, but I would probably be wary of admitting to participate in the illegal book piracy scene unless your pseudonymity is really tight. Thanks a lot for helping out, though.

I did something similar: OCR'd a book, then translated it to another language to make it accessible for a friend.

However, I also contacted the author, and sent them a copy! At first, they were furious - I'd even translated ('pirated') the copyright notice...

But things settled down, and now I'm working with the author on translations to various other languages. It's given the book a whole new audience!

Re: Pirate Library Mirror: Preserving 7TB of books (that are not in Libgen)

#216

Earlier quoted context omitted.

Google Books, like so many Google projects, had a dual purpose. Making books accessible is noble and on-mission. But more importantly natural language models can be trained on the scanned corpus. The same was true of the original GOOG 411, which provided a free service, but was really put in place to train up their voice recognition projects. This is a long running strategy of Google, and it's a shrewd one. The main…

That can't be true, Google Books was 15 years prior to the advent of large language models. Until 2020 nobody could train on such a large collection. I think Google initially wanted to augment the web results with a large book collection to get "all the world information and make it searchable", same with Google News.

Before there were neural language models, there were n-gram models, skip n-gram models, latent dirichlet allocation models, a whole zoo of non-neural machine learning models. Google used some of them to power old Google Translate, which was still a very impressive piece of technology.

Re: Pirate Library Mirror: Preserving 7TB of books (that are not in Libgen)

#217
post #210

Excellent. I hope more and more people start to see how absurd and evil the concept of "intellectual property" is. It should be totally rejected and any form of keeping useful information for yourself should be shunned and tabooized. In todays world, many have been programmed to believe that the would could not exits without such immoral restrictions, which is horrible.

I've just spent 2 years writing something which ain't got anything else like it. It was technically pretty difficult and needed a lot of background knowledge.

Should I be disallowed to commercialise it?

I partly get where you stand but if I was in a society that you seem to endorse my first question would be, other than for the love of doing it, why sink so much effort into a thing only to get nothing back. It almost is the opposite of a meritocracy.

Re: Pirate Library Mirror: Preserving 7TB of books (that are not in Libgen)

#219
post #124

Again, it's just insane to me that we don't even much have a meaningful discussion of: "Hey, wait, literally everyone could have the entire library of Alexandria in their house for a couple hundred bucks per person. Like, all the knowledge ever. Maybe that should be considered the good default of things. At least one in every town that everyone could use, for free, forever, without restriction to ANY of the knowledge…

As a kid, my parents had a full copy of the Library of Alexandria on the coffee table. Everyone else just called it an ashtray. (Sorry, too soon?)

:'(

Re: Pirate Library Mirror: Preserving 7TB of books (that are not in Libgen)

#220
post #124

Again, it's just insane to me that we don't even much have a meaningful discussion of: "Hey, wait, literally everyone could have the entire library of Alexandria in their house for a couple hundred bucks per person. Like, all the knowledge ever. Maybe that should be considered the good default of things. At least one in every town that everyone could use, for free, forever, without restriction to ANY of the knowledge…

I think it's particularly insane that (ignoring scihub, which still faces legal battles and is at risk of being our next Library of Alexandria) the world's scientific knowledge is largely behind paywalls and inaccessible to most of humanity, even those millions whose taxes funded it.
Post reply on HN