Live data from Hacker News

Pirate Library Mirror: Preserving 7TB of books (that are not in Libgen)

pilimi.org

221–230 of 438 posts

Re: Pirate Library Mirror: Preserving 7TB of books (that are not in Libgen)

#221
post #210

Excellent. I hope more and more people start to see how absurd and evil the concept of "intellectual property" is. It should be totally rejected and any form of keeping useful information for yourself should be shunned and tabooized. In todays world, many have been programmed to believe that the would could not exits without such immoral restrictions, which is horrible.

There should be a fine line in intellectual property rights. I see where you are coming from - quite often intellectual property is used a a moat to protect insane revenues and, as a repercussion, delay or slowdown our progress as humanity. But it is also use to protect unique creator revenue and encourage to create more. If you ask where the fine line should be I have no immediate answer, but abolishing intellectual…

I think that's pretty much the case for shortening time frame of IP rights. Which means, they shouldn't be treated as property has been traditionally treated. Although as a socialist, I think perhaps we shouldn't have time unlimited property rights (above certain reasonable boundary, say $10M) in general.

Re: Pirate Library Mirror: Preserving 7TB of books (that are not in Libgen)

#222

Earlier quoted context omitted.

That can't be true, Google Books was 15 years prior to the advent of large language models. Until 2020 nobody could train on such a large collection. I think Google initially wanted to augment the web results with a large book collection to get "all the world information and make it searchable", same with Google News.

> That can't be true, Google Books was 15 years prior to the advent of large language models. Until 2020 nobody could train on such a large collection. ... pull the other one. Okay, your statement could be true depending on what you mean by "large". But what makes you think that companies like Google haven't been working on language models without releasing them and/or without discussing them publicly? There's an adv…

The gap between research and publication is real but its length is less than one year. I don't think even Google has the resources to do it secretly. Who would work on it and how could they have kept the secret so tight? AI researchers want to publish, especially the best ones, it's essential for their careers. Someone else could plant the flag on their discovery and claim the fame.

Re: Pirate Library Mirror: Preserving 7TB of books (that are not in Libgen)

#223
post #182

Earlier quoted context omitted.

That's already the case minus the last ~70 years or so. The overwhelming majority of our knowledge is in the public domain, in particular cultural artifacts. It's a nice sentiment but like, people can already go to gutenberg.org and download pretty much most important works of literature in existence and most books have like 5k downloads so there's that.

Project Gutenberg is missing a lot of content that is public domain, and the oldest entries are often pretty poor in quality (and some of the oldest entries aren't lucky enough to get redone like Carroll's work has been).

Which just goes to prove parent's point. As much as it might seem otherwise, IPR restrictions are not the main bottleneck to widespread availability of content (at least book-like content); actually making the content available is far more important!

(Also, keep in mind that content is now entering the public domain every year, and projects like PG are nowhere close to keeping up with that flow of newly-unrestricted stuff. So this dynamic is becoming more extreme over time, not less.)

Re: Pirate Library Mirror: Preserving 7TB of books (that are not in Libgen)

#224
post #210

Excellent. I hope more and more people start to see how absurd and evil the concept of "intellectual property" is. It should be totally rejected and any form of keeping useful information for yourself should be shunned and tabooized. In todays world, many have been programmed to believe that the would could not exits without such immoral restrictions, which is horrible.

I've just spent 2 years writing something which ain't got anything else like it. It was technically pretty difficult and needed a lot of background knowledge. Should I be disallowed to commercialise it? I partly get where you stand but if I was in a society that you seem to endorse my first question would be, other than for the love of doing it, why sink so much effort into a thing only to get nothing back. It almost…

If you as a private person own a patent, you are losing it anyways...because you cannot fight some mega corporation in court to defend it. It's too expensive, and the biggest corps just take what they want due to having more financial resources.

Note that if you don't defend it in court, our justice system thinks it is less valueable to protect. Which in itself is kind of ridiculous.

Also, right to commercialization has nothing to do with intellectual property.

Re: Pirate Library Mirror: Preserving 7TB of books (that are not in Libgen)

#225
post #211

Earlier quoted context omitted.

There are some language models like DeepMind RETRO that can make use of a 1TB text collection after the model is trained. The idea is to chunk up the text and make the blocks searchable with a neural embedding index. When you ask a question to the model, it first searches for relevant information and then adds it to the prompt. The result - you can get GPT-3 quality on a 25x smaller model. That means you could have y…

The interesting aspect of GPT-3+ is their reasoning like behavior such as Chains of thought or "step by step" generation. There's no evidence RETRO as is has sufficient capacity for the actually interesting behaviors.

I agree, that's why I asked the authors the same question : https://twitter.com/visarga/status/1543826941052166144

Re: Pirate Library Mirror: Preserving 7TB of books (that are not in Libgen)

#226

I wonder if there are any search engines dedicated to indexing these kinds of libraries. I know there's a decent one just for scihub, but it would be awesome if I could do a Google-style search that returned the contents of books, magazines and journal articles instead of just websites.

Book metadata is widely available via sites like e.g. Open Library. With good metadata, full text search is not as relevant.

That's false. I often search for specific citation.

Re: Pirate Library Mirror: Preserving 7TB of books (that are not in Libgen)

#227
post #221

Earlier quoted context omitted.

There should be a fine line in intellectual property rights. I see where you are coming from - quite often intellectual property is used a a moat to protect insane revenues and, as a repercussion, delay or slowdown our progress as humanity. But it is also use to protect unique creator revenue and encourage to create more. If you ask where the fine line should be I have no immediate answer, but abolishing intellectual…

I think that's pretty much the case for shortening time frame of IP rights. Which means, they shouldn't be treated as property has been traditionally treated. Although as a socialist, I think perhaps we shouldn't have time unlimited property rights (above certain reasonable boundary, say $10M) in general.

Interesting, but being very opposite of socialist myself I am wholeheartedly with you here on limiting timeframe of IP. It partially solves the problem.

And this arguably should be extended to tangible assets as well - I like Singapore model where housing property is sold for specific timeframe. It simplifies a lot of redevelopment.

Re: Pirate Library Mirror: Preserving 7TB of books (that are not in Libgen)

#228

It's a shame that this would have been a textbook case for using IPFS and yet that wasn't the default. Books are naturally immutable, and could be structured into sub-categories whilst enjoying the benefits of deduplication.

IPFS doesn't really work well for this because you'd never know if the peer hosting the last subset of some books went offline (and you'd lose those until someone who had them came online again). I want a slightly different system, which I've posted about here: https://news.ycombinator.com/item?id=31972252

This is my point about the shame it isn't. It seems obvious that ipfs would have both privacy and a self balancing way to pin a partial set of data. But no, which makes it unsuitable.

Re: Pirate Library Mirror: Preserving 7TB of books (that are not in Libgen)

#229
post #210

Excellent. I hope more and more people start to see how absurd and evil the concept of "intellectual property" is. It should be totally rejected and any form of keeping useful information for yourself should be shunned and tabooized. In todays world, many have been programmed to believe that the would could not exits without such immoral restrictions, which is horrible.

There should be a fine line in intellectual property rights. I see where you are coming from - quite often intellectual property is used a a moat to protect insane revenues and, as a repercussion, delay or slowdown our progress as humanity. But it is also use to protect unique creator revenue and encourage to create more. If you ask where the fine line should be I have no immediate answer, but abolishing intellectual…

> But it is also use to protect unique creator revenue and encourage to create more.

This thinking is an artifact of an economic system so dependent on scarcity for its motivation that it is now generating most of the scarcity in the world.

We now have the technology for creative implementations of "From each according to its ability, to each according to its needs". Just keep track of how much each thing is used, and reward creators from a corporate-tax-funded pool. Every for-profit entity contributes proportionally to its profit, and can use any idea for free.

Re: Pirate Library Mirror: Preserving 7TB of books (that are not in Libgen)

#230
post #77

Earlier quoted context omitted.

What kind of fairness can there be in charging for stolen books? I believe in free access to education, but charging for these books they have no rights to is a whole other thing.

I don't quite agree. I mean, they provide useful service, and it costs money to run it. It's ok that they earn (even if it's actually making a profit, not just covering the costs). That being said, 10 downloads/day feels a bit restrictive to me. I'd get if it was 100, or 50, heck, maybe even 20. I mean, I don't appreciate that it's not mirrorable in the first place, but maybe they cannot afford it, I don't know... Bu…

Whenever you reach the limit, you can always copy the ISBN and search it on Libgen.

8 times outta 10, the book would be there.

Post reply on HN