Live data from Hacker News

Pirate Library Mirror: Preserving 7TB of books (that are not in Libgen)

pilimi.org

191–200 of 438 posts

Re: Pirate Library Mirror: Preserving 7TB of books (that are not in Libgen)

#191
post #124

Again, it's just insane to me that we don't even much have a meaningful discussion of: "Hey, wait, literally everyone could have the entire library of Alexandria in their house for a couple hundred bucks per person. Like, all the knowledge ever. Maybe that should be considered the good default of things. At least one in every town that everyone could use, for free, forever, without restriction to ANY of the knowledge…

That's already the case minus the last ~70 years or so. The overwhelming majority of our knowledge is in the public domain, in particular cultural artifacts. It's a nice sentiment but like, people can already go to gutenberg.org and download pretty much most important works of literature in existence and most books have like 5k downloads so there's that.

let's not forget what aaron swartz died for, though. quite a bit of science from publicly funded studies is behind paywalled publications.

Re: Pirate Library Mirror: Preserving 7TB of books (that are not in Libgen)

#192
post #124

Again, it's just insane to me that we don't even much have a meaningful discussion of: "Hey, wait, literally everyone could have the entire library of Alexandria in their house for a couple hundred bucks per person. Like, all the knowledge ever. Maybe that should be considered the good default of things. At least one in every town that everyone could use, for free, forever, without restriction to ANY of the knowledge…

A tablet with WiFi/cellular for every person would go a long way.

Re: Pirate Library Mirror: Preserving 7TB of books (that are not in Libgen)

#193
post #124

Again, it's just insane to me that we don't even much have a meaningful discussion of: "Hey, wait, literally everyone could have the entire library of Alexandria in their house for a couple hundred bucks per person. Like, all the knowledge ever. Maybe that should be considered the good default of things. At least one in every town that everyone could use, for free, forever, without restriction to ANY of the knowledge…

> "Hey, wait, literally everyone could have the entire library of Alexandria in their house for a couple hundred bucks per person. Like, all the knowledge ever. Maybe that should be considered the good default of things.

There are projects out there that lean in the direction of offline viewing of lots of content, for example, having an offline backup of Wikipedia, such as:

  - https://wiki.kiwix.org/wiki/Main_Page
  - http://xowa.org/ (HTTPS seems not to work though)
I just wish that the process of actually accessing the data was a little bit more straightforward: https://dumps.wikimedia.org/backup-index.html (given how many different files there are to choose from, you'll probably want to read a tutorial or two)

That said, while text is perfectly doable, things do tend to get more difficult if you also want images or videos, because those do take up a lot of space.

Re: Pirate Library Mirror: Preserving 7TB of books (that are not in Libgen)

#195

Earlier quoted context omitted.

Google Books, like so many Google projects, had a dual purpose. Making books accessible is noble and on-mission. But more importantly natural language models can be trained on the scanned corpus. The same was true of the original GOOG 411, which provided a free service, but was really put in place to train up their voice recognition projects. This is a long running strategy of Google, and it's a shrewd one. The main…

That can't be true, Google Books was 15 years prior to the advent of large language models. Until 2020 nobody could train on such a large collection. I think Google initially wanted to augment the web results with a large book collection to get "all the world information and make it searchable", same with Google News.

> That can't be true, Google Books was 15 years prior to the advent of large language models. Until 2020 nobody could train on such a large collection.

... pull the other one.

Okay, your statement could be true depending on what you mean by "large". But what makes you think that companies like Google haven't been working on language models without releasing them and/or without discussing them publicly? There's an advantage to be had by keeping corporate secrets.

Re: Pirate Library Mirror: Preserving 7TB of books (that are not in Libgen)

#196

Earlier quoted context omitted.

Ah hello Walter. I recognize your username as you and I have disagreed on this before. I’m not saying people shouldn’t have the right to own property. But that a good way of organizing society is collective ownership of the means of production. If you are part owner in something with shares and a contract, that obviously still relies on property rights. That is how the stock market works after all. EDIT: I am basical…

The US is a free country. You can form a voluntary collective any time you like.

To be perfectly honest this feels like a knee jerk response that fails to engage with what I am saying. Nowhere did I say I am being prevented from forming a voluntary collective. But obviously if I believe we should form an economy comprised of many collectives then I need to discuss this idea with other people! I literally cannot form a collective by myself. For my part I am working on a career trajectory that will allow me to co-found a collective that owns machines which produce free hot meals for people. But that is a few years away. Still I find it interesting that you want to argue without really thinking about what I am saying. Obviously forming collectives necessarily involves discussing these ideas openly with others.

Re: Pirate Library Mirror: Preserving 7TB of books (that are not in Libgen)

#197

Earlier quoted context omitted.

Ah hello Walter. I recognize your username as you and I have disagreed on this before. I’m not saying people shouldn’t have the right to own property. But that a good way of organizing society is collective ownership of the means of production. If you are part owner in something with shares and a contract, that obviously still relies on property rights. That is how the stock market works after all. EDIT: I am basical…

Changing norms is a nice idea, but you'll have to lead by example I believe. That said, I think privately- and collectively-owned enterprises can coexist (and compete!) just fine, and in a truly free market we'd see a healthy amount of both.

I am working on leading by example, though putting the pieces together for my career will take some time. I have so far already been working on this for a few years. I am spinning up a non profit open source project to design solar powered farming robots and I run two YouTube channels to promote my economic ideas and the farming robot project. Still, I do like to discuss the merits of the base philosophy here.

Re: Pirate Library Mirror: Preserving 7TB of books (that are not in Libgen)

#198

Earlier quoted context omitted.

> https won't keep your ISP from knowing you visited the site If you use DoH, yes it does. Unless I'm mistaken. They only know the IP address of the remote server.

> They only know the IP address of the remote server. It's the internet. Everyone can scrape links and measure/correlate which assets were on them to correlate likely visited websites. Especially if every web page these days is pretty unique in terms of what kind of assets (network streams) with what kind of byte size were loaded at which point in the document loading timeline. Now include the TLS fingerprint of your…

> HTTP needs an upgrade with scattering and rerouting on the fly, otherwise these deanonymization techniques can never be fixed.

Isn't that e.g. TOR's job? Doesn't belong in HTTP.

Re: Pirate Library Mirror: Preserving 7TB of books (that are not in Libgen)

#199

I see things like this, and I wonder why the following software doesn't exist: I want a piece of software to which I can add a collection of files, say multiple TB. The software will then behave a bit like a BitTorrent tracker, and know which peer has which files. A peer joining this swarm will be able to say "I want to donate X GB of space", and the tracker would tell it "OK, then download and seed these files, whic…

That's literally how most of Japanese P2P software work.

For example, Perfect Dark, Winny, and to less extent, Share (which is more similar to eDonkey/eMule).

Post reply on HN