Live data from Hacker News

Pirate Library Mirror: Preserving 7TB of books (that are not in Libgen)

pilimi.org

401–410 of 438 posts

Re: Pirate Library Mirror: Preserving 7TB of books (that are not in Libgen)

#401
post #311

Earlier quoted context omitted.

> something in the realm of 20 years should be enough. God no. Right now, we'd be getting remakes from every piece of pop-culture that was semi-popular in the 80s-2002 time frame. Not just movies, but TV series, books, theatre, musicals, ... Sure, copyright should be shortened, but I don't begrudge (eg.) a one hit winner making money off their hit decades layer. Life of artist is reasonable, I think. They take a gamb…

I think 30 years is a decent base in my mind. Let it be renewed a couple times for an extra years if the owners think it’s worth it. That way most stuff flows into the public domain but some works can keep creating value for their creator.

Upvote for 30 years. Try it for a while, if creative industries aren't meaningfully damaged by the reduced IP rights, then maybe even take it down lower to like 25 or even 20 years.

Re: Pirate Library Mirror: Preserving 7TB of books (that are not in Libgen)

#402

Earlier quoted context omitted.

Cinema rips were pretty bad even before cinema's started filming the audience.. Hardly worth watching. Other than that, all DRM does it makes it harder, not impossible, to copy.

«Cinema's started filming the audience»?! You sign some consent form? In which country?

You're in their building, so I imagine your consent is stated as part of the ticket verbiage or in small print when you buy it. Regardless, closed-circuit cameras are pretty ubiquitous in private businesses, and have been for decades.

Re: Pirate Library Mirror: Preserving 7TB of books (that are not in Libgen)

#403
post #124

Again, it's just insane to me that we don't even much have a meaningful discussion of: "Hey, wait, literally everyone could have the entire library of Alexandria in their house for a couple hundred bucks per person. Like, all the knowledge ever. Maybe that should be considered the good default of things. At least one in every town that everyone could use, for free, forever, without restriction to ANY of the knowledge…

OK, what's the napkin math?

I've specced this out a few times, it's an interesting exercise, and it depends on what "all human knowledge" is.

From the perspective of how much you individually could read in a lifetime: that's about 4,000 weeks (to roughly age 80-ish, and presuming you're not already reading at birth). Multiply that by the number of books you plan on reading a week. You'll probably read fewer than 4,000 books over your entire life. Many people read few if any books after graduating secondary school, and even a highly-motivated reader might be challenged to crack, say, 40,000 (ten per week for life).

The US Library of Congress has the world's largest book collection at about 40 million distinct titles.[1] There are another 130 million total catalogued items, including photographs, films, audio recordings, maps, pamphlets, and other items. Not all are textual. I'll stick to books.

At 4,000 per lifetime, you'd need 10,000 people simply to read all of the LoC's collection at a reasonable pace.

(By comparison, at the turn of the 20th century, the LoC's annual report to Congress noted that its cataloguing department could handle about 3,000 books per cataloguer per year, or a pace of about 60/week or 15/day. Over a 40-year career, a cataloguer might handle, if not read completely, 120,000 books.)

As a rough rule of thumb, a digitised ebook in PDF format runs about 5 MB.

Every single book in the Library of Congress in digital format would occupy about 200 TB of storage.

Current disk prices are running about $5 -- $20 / TB.

200 TB in raw disk would set you back $1,000 -- $4,000. Figure 2-4x multiplier for a disk storage system all told, and it's still roughly $4k -- $16k to have as local storage every last book in the US Library of Congress. That's well within scope of a moderately wealthy US household budget, and would be reasonable to consider for a small-town city library.

That's the technical storage cost, obviously not the rights or aquisition costs for the materials.

The Library of Congress's budget is about $800 million/yr.

Those 4,000 books you might read in a lifetime? They'd fit on about 20 GB worth of disk. If you're a 10x reader, 200 GB, and a truly dedicated 100x reader, about 2 TB. Those are well within the range of present desktop / laptop disk allocations, and represent a few tens of dollars of storage expense.

The Internet Archive computes $2/GB for storage in perpetuity. That's roughly 400 books worth of storage.

But a small household NAS and a very modest server platform (most routers can effectively operate as media servers) could rival a mid-sized city library for about $100 or less in actual hardware outlay. Those prices are falling by half about every 18 to 36 months, as they have been for decades.

Books are not all published content, and all published content is not all human knowledge. But standard published books are a good proxy for total cultural knowledge, and in all likelihood, then some.

________________________________

Notes:

1. See: https://www.loc.gov/about/general-information/ That includes about 25 million books in the main collection, and 15 million in "nonclassified print collections, including books in large type and raised characters, incunabula (books printed before 1501), monographs and serials, music, bound newspapers, pamphlets, technical reports and other printed material". I suspect you could roughly halve my estimates above based on 25 million vs. 40 million volumes, as many of the nonclassified works may duplicate the main collection.

Re: Pirate Library Mirror: Preserving 7TB of books (that are not in Libgen)

#404
post #377

Earlier quoted context omitted.

You can always implement QoS or something like FQ_Codel.

I'm a senior software developer with experience in developing communications protocols and networking software. I haven't been able to set up QoS properly ever, especially in a way that addresses all my needs. Expecting end users to do it is way beyond the realm of possibility IMHO.

It is a huge pain in the ass to get it working right I will admit. Took me hours of reading docs and tweaking parameters. I'm still not sure I understand all of it, but I managed to get it working such that it at least meets my needs.

Re: Pirate Library Mirror: Preserving 7TB of books (that are not in Libgen)

#405
post #124

Again, it's just insane to me that we don't even much have a meaningful discussion of: "Hey, wait, literally everyone could have the entire library of Alexandria in their house for a couple hundred bucks per person. Like, all the knowledge ever. Maybe that should be considered the good default of things. At least one in every town that everyone could use, for free, forever, without restriction to ANY of the knowledge…

That's already the case minus the last ~70 years or so. The overwhelming majority of our knowledge is in the public domain, in particular cultural artifacts. It's a nice sentiment but like, people can already go to gutenberg.org and download pretty much most important works of literature in existence and most books have like 5k downloads so there's that.

It's actually fairly unlikely that the bulk of published knowledge is in the public domain.

By copyright expiry, public domain in the US begins in 1927. Later works may be in the public domain, but all works published prior to 1927 are in the public domain in the United States.

(This may not be the case in other countries.)

There were not many published books prior to the invention of the printing press, and many of those didn't survive. The total number of books (not individual titles but actualy bound volumes entirely) in Western Europe as of 1400 may have been as few as 50,000.

By 1800, about 1 million titles had been printed.

Over the course of the 18th century, presses became vastly faster, as they evolved from hand-operated wooden screw-press to iron-frames to steam and electric-powered rotary and ultimately web presses. Paper became much cheaper (and less durable --- a factor commented on at length in the Librarian of Congress's annual reports to Congress in the late 19th century). Literacy exploded from ~25% to 95%+ over the 19th century (and probably accounted for numerous revolutions and political upheavals).

Through much of the 20th century, certainly by 1950, US publishers were issuing about 300,000 new titles per year, a rate which state remarkably constant through the early 21st century. By the aughts, "nontraditional" self-publishing (a/k/a "vanity press") was nearing or exceeding 1 million titles per year more than had been published through all time to 1800.

Reports that all recorded data was doubling every few years date to at least the 1960s. That would mean that in any two year period ... half of all recorded information was less than two years old.

The catch is that not all recorded data is published. So I'm not sure what the time-distribution of all publishing looks like. But I'm pretty confident it's skewed far more recently than 1927. And would thus tend to be copyrighted rather than uncopyrighted.

If you want to measure works by significance, you might make a different argument --- there are many great works of literature, philosophy, history, and religion which were first published before 1927. But ranking and tabulating these is more challenging than a simple enumeration.

Re: Pirate Library Mirror: Preserving 7TB of books (that are not in Libgen)

#406
post #185

Earlier quoted context omitted.

That's already the case minus the last ~70 years or so. The overwhelming majority of our knowledge is in the public domain, in particular cultural artifacts. It's a nice sentiment but like, people can already go to gutenberg.org and download pretty much most important works of literature in existence and most books have like 5k downloads so there's that.

Hasn't been more knowledge published in the last 70 years than during all the times before? More than 2 million new books get published every year.

Pretty much certainly: https://news.ycombinator.com/item?id=31984080

Re: Pirate Library Mirror: Preserving 7TB of books (that are not in Libgen)

#407

Earlier quoted context omitted.

"The African economy will start booming once they get UNLIMITED access to those 1800s manuals on measuring the Ether and where the best seal clubbing sites are in Alaska." Nobody ever thinks this through. Books are basically worthless at this point

then maybe we stop letting people push the copyright periods tentatively so more relevant things come into the public domain as well. this is a nothing argument.

There's a compelling argument that Germany's lax attitudes toward copyright in the 19th century helped propell that nation[1] to a dominant position in science and technology.

Germany now is a leading copyright-maximalist power, unfortunately.

________________________________

Notes:

1. I say nation and not country as the German language and cultural nationality pre-existed the German state in 1871.

Re: Pirate Library Mirror: Preserving 7TB of books (that are not in Libgen)

#408

Earlier quoted context omitted.

> Gutenberg was blocked in Italy in early 2020. I don't know if this block still persists. afaict, it resolves and loads fine here. Now on mobile tethering via the Iliad (.it) local operator. When I get home I'll also check it on broadband.

I have information that the abysmal filth is still in place: "Sito sotto sequestro". I wonder how your provider overrides it. Is it too much of a wet dream to think that some operators (behind the DNS) responded to the requests of the vacuously appointed low scum* with disdainful neglect? In terms of "If you really had any credential for existence but some forms of physical expression, we would say that you must be j…

Just tried it at home and loads fine from broadband too. That "Sito sotto sequestro" literally means "site has been seized", which of course isn't the case here; that page might indicate that the carrier has been forced to resolve that domain elsewhere, therefore all it needs is using an external name server to get around the block. I'm currently using one of Cloudflare DNS addresses and it loads just fine.

I don't use the (ex) national provider (Telecom Italia / Tim) however. Crappy service aside, they're well known for jumping every time the government tells them to, and happily filter a lot of stuff.

Re: Pirate Library Mirror: Preserving 7TB of books (that are not in Libgen)

#409

Earlier quoted context omitted.

maybe it has something to do with the commercial nature of one of those two things.

In order to run Codex you need 4-8 large GPUs connected in the same box or cabinet. Such computers cost about $100k to buy. Then they use about 5..10kwh to run, that adds up to 50MW/year. The Codex model needs to be very low latency, means more expensive. This is just for one single replica of the model, you need many to serve all the requests. Not to mention the cost of development - for that they needed thousands o…

I'm sorry but can you run a mirror for possibly the largest collection of written knowledge ever assembled in your home? What exactly is the point of this exercise?

Re: Pirate Library Mirror: Preserving 7TB of books (that are not in Libgen)

#410

Earlier quoted context omitted.

> That can't be true, Google Books was 15 years prior to the advent of large language models. Until 2020 nobody could train on such a large collection. ... pull the other one. Okay, your statement could be true depending on what you mean by "large". But what makes you think that companies like Google haven't been working on language models without releasing them and/or without discussing them publicly? There's an adv…

The gap between research and publication is real but its length is less than one year. I don't think even Google has the resources to do it secretly. Who would work on it and how could they have kept the secret so tight? AI researchers want to publish, especially the best ones, it's essential for their careers. Someone else could plant the flag on their discovery and claim the fame.

For a time while I was working at Google I nursed the idea of transferring to the natural language processing / language modeling areas, but I only have a bachelors in linguistics whereas it seems like they strongly prefer PhDs. Linguistics PhDs can pretty much teach linguistics or go into another field. It would not be hard for Google to find a couple hundred of them and entice them to join up.
Post reply on HN