Fuck that site. Offers people links to free PDF downloads of my book that I worked on for 32 years and finally got published by Pantheon Books in 2017. I didn't work all that fucking time for criminals like these to just break copyright law and make the book available for free. Fuck Anna's Archive, and I hope they go down in legal flames ASAP.
Anna's Archive: An Update from the Team
351–360 of 560 posts
Re: Anna's Archive: An Update from the Team
#352remember guys, it's not pirating, it's gathering date from AI model training purposes. Perfectly legal.
Re: Anna's Archive: An Update from the Team
#353Earlier quoted context omitted.
Yes, but whereas libraries need to buy more copies of books that lots of people check out, Anna's archive only ever needs one. Not exactly sustainable for the author. As I said, I loved the Internet Archive's approach to this! That's very much not what Anna's archive is doing.
Still, libraries buy what, maybe 5 copies of a mildly popular book. I don't think that would be sustainable either if that was the only books sold.
Re: Anna's Archive: An Update from the Team
#354Earlier quoted context omitted.
Kind of... the fact that they have the actual data behind a "soft" paywall (waiting times and terribly slow transfers otherwise) makes me a bit skeptic of their "goodwill".
Bandwidth isn’t free of charge
Re: Anna's Archive: An Update from the Team
#355Earlier quoted context omitted.
Interesting to see that sci-hub is about 90TB and libgen-non-fiction is 77.5TB. To me, these are the two archives that really need protecting because this is the bulk of scientific knowledge - papers and textbooks. I keep about 16TB of personal storage space in a home server (spread over 4 spinning disks). The idea of expanding to ~200 TB however seems... intimidating. You're looking at ~qty 12 16TB disks (not counti…
A lot of these are (relatively large) pdfs, right? I wonder how much space it is as highly compressed, deduplicated, plain text files. Does the sum of human scientific knowledge fit on a large hard drive?
The problem is it's not all text, you need the images, the plots, etc, and smartly, interstitially compressing the old stuff is still a very difficult problem even in this age of AI.
I have an archive of about 8TB of mechanical and aerospace papers dating back to the 1930s, and the biggest of them are usually scanned in documents, especially stuff from the 1960s and 70s, that have lots of charts and tables that take up a considerable amount of space, even in black and white only, due to how badly old scans compress (noise on paper prints, scanned in, just doesn't compress). Also many of those journals have the text compressed well, but they have a single, color, HUGE cover image as the first page of the PDF, that turns the PDF from 2MB into 20MB. Things like that could, maybe, be omitted to save space...
But as time goes on I start to become more against space-saving via truncation of those kind of scanned documents. My reasoning is that storage is getting cheaper and cheaper, and at some point the cost to store and retrieve those 80-90MB PDF's that are essentially total page by page image scans is going to be completely negligible. And I think you lose something be taking those papers and taking the covers out, or OCR'ing the typed pages and re-typesetting them to unicode (de-rasterize the scan), even when done perfectly (and when not done perfectly, you get horrible mistakes in things like equations, especially). I think we need to preserve everything to a quality level that is nearly as high as can be.
Re: Anna's Archive: An Update from the Team
#356Earlier quoted context omitted.
Still, libraries buy what, maybe 5 copies of a mildly popular book. I don't think that would be sustainable either if that was the only books sold.
Libraries have to replace paperback books after ~20 checkouts on average. (This number is from memory but I'm quite sure it's in this range.) Hardcover books last a bit longer but of course are also more expensive. I agree the industry would have a hard time surviving off library sales alone, in the same way that most businesses rely on multiple revenue streams to make ends meet, but I think library revenue is much m…
As for anecdota, I have more than once borrowed a library book and then purchased a copy so I could read it again or to finish it if demand is strong enough that I would have to wait weeks or months to be able to borrow it again.
Re: Anna's Archive: An Update from the Team
#357Earlier quoted context omitted.
I would not buy a book after downloading it from Anna's archive. But that's the wrong question in my opinion. You should be asking why aren't most books available in a DRM free format? The main reason to download "pirated" books is that they get rid of all annoying barriers that exist in "legitimate" copies. It's a better product .
> You should be asking why aren't most books available in a DRM free format? Because most people don't care! I wish they did, because I'm like you, I do care about owning DRM free media! I buy videos game from GOG wherever possible, and audiobooks from a combination of downpour.com and libro.fm. Guess what most people do? They buy games on Steam and audiobooks on Audible. Audible is the one that really breaks my hear…
Just had a browse of Downpour. They say that it's mostly DRM-free. I don't get it. How come the rights holders don't complain? My experience of DRM-free e-books is that the available titles are, let's say, nothing I would want to read. And audiobooks have higher production value because of the voice acting. What A-list authors are narrating their own books and then allowing them to be sold DRM-free?
Re: Anna's Archive: An Update from the Team
#358Earlier quoted context omitted.
> Authors and rights holders are supposed to just take it? If it's judged as fair use, then yes. And then it's not flouting anything. Remember the whole point of fair use is to benefit society by allowing reuse of material in ways that don't directly copy large portions of the material verbatim. For example, nonfiction authors already "just take it" when reviews describe the main points of their book without paying t…
If I was a writer, I'd consider publishing my works under a license that explicitly bans AI training. What happens when those works inevitably get ingested by an LLM?
Your license can only operate with what copyright allows you to withhold initially.
A license that banned AI training cannot be enforced. It is meaningless. The same way you can't write a book with a license that readers are not allowed to write reviews of it.
Fair use cannot be restricted by license like that.
(You can engage in individual contacts with people, with terms like NDA's work, but those actually have to be signed and stuff, and you can't do it with public information like published writing.)
Re: Anna's Archive: An Update from the Team
#359Earlier quoted context omitted.
The idea that any widely transmitted identifiers' confidentiality should be its primary method of security is asinine. The failure of any exploit regarding SSNs or the like is not on the offending party, but on each using party's failure to implement even a modicum of actual security.
FYI calling something "asinine" is not an argument.
Re: Anna's Archive: An Update from the Team
#360Know am going to be downvoted into oblivion, but as a composer, can see it from the side of creators. Yeah, making their products free is starving these industries. For instance, in music, there is already very little money in music (think about how many musicians you personally know who can make a living off of music, besides being a music teacher). And, the music industry is still not even the same size as it was i…
Both producers and consumers of media are in the same boat of barely surviving. Maybe we can work with each other instead of against each other? :)