Live data from Hacker News

Google Books (or similar) all book scans – $200k bounty (2025)

software.annas-archive.gl

301–310 of 368 posts

Re: Google Books (or similar) all book scans – $200k bounty (2025)

#301

Earlier quoted context omitted.

The reason I think it is hard in the US is because there is a very strong "work or die" ethic in the US. Everything is driven by money. Even basic things like healthcare are driven by money. Your life after retirement is determined by how much money you accumulated. The word-association between "poor" and "lazy" is strong. Taxation should be light. Each man should keep what he accumulates. BI by contrast values peopl…

Top 10% already pays 70.5% of all federal income taxes. US high income payers are already taxed to pay the poor. Almost half of US federal tax payers pay 0 income tax.

Another way of looking at that is that the people who have been able to accumulate excessive gains from the capitalist system are forced to pay some of that back, to maintain the system that enriched them and those whose labor they profited from.

That seems like a screwy but ultimately more than fair deal for the top 10%.

Especially when the alternative is pitchforks and torches.

Re: Google Books (or similar) all book scans – $200k bounty (2025)

#302
post #125

Gemini should be trained on those books already, so in theory it could regurgitate some verbatim fragments (as NYT lawsuit against OpenAI showed some time ago).

Gemini, gpt and fable are actually very good compressions of internet content. But is lossless compression as in they kept the most important part (for them to fulfill the next token task) and found a way to mimic the rest.

I think you meant lossy compression and not lossless. I'm not suggesting this as a method to extract those books from the models, which by their nature are not databases. Just commenting on the somewhat surprising fact that the bigger the model the more likely it is to produce some (short) excerpts of the original training material

Re: Google Books (or similar) all book scans – $200k bounty (2025)

#303
post #299

Earlier quoted context omitted.

Not so. Stallman created copyleft licenses as a defense against the current implementation of copyright. Copyleft uses the existing system of copyright to protect authors of free software from people who want to use copyright to restrict distribution. It wouldn't be necessary if copyright didn't exist.

Stallman wanted to protect the right to fix bugs, he was not against paying for goods and services.

Sure, but an idea is not a (physical) good, nor is it a service. Coming up with an idea or writing a book is a service and should be paid for (probably by commission), but (and Stallman would agree) the idea or book itself should be free.

Re: Google Books (or similar) all book scans – $200k bounty (2025)

#304
post #301

Earlier quoted context omitted.

Top 10% already pays 70.5% of all federal income taxes. US high income payers are already taxed to pay the poor. Almost half of US federal tax payers pay 0 income tax.

Another way of looking at that is that the people who have been able to accumulate excessive gains from the capitalist system are forced to pay some of that back, to maintain the system that enriched them and those whose labor they profited from. That seems like a screwy but ultimately more than fair deal for the top 10%. Especially when the alternative is pitchforks and torches.

Before you go to the capitalist in their home and torch it or beat them in front of their family, try striking for a while. Usually they soften rather quickly when labour is collectively withheld.

Re: Google Books (or similar) all book scans – $200k bounty (2025)

#305

Earlier quoted context omitted.

If this is not bad faith argument, then I don't what is. When someone is violating an OSS licence, they are doing it for commercial gains and monetary profit. Nobody is angry at someone using FOSS software for himself with no money getting involved. As opposed to that, books, movies are pirated for personal consumption. Not monetary gains. If someone bought a $30 book, and then ran a BaaS with millions of VC money in…

Anthropic, OpenAI and Meta: yeah totally, personal consumption only

Altman et al take knowledge and lock it a way. The add restraints, borders, limitations of things they got for free. They stole freedoms.

Re: Google Books (or similar) all book scans – $200k bounty (2025)

#306
post #304
post #301

Earlier quoted context omitted.

Another way of looking at that is that the people who have been able to accumulate excessive gains from the capitalist system are forced to pay some of that back, to maintain the system that enriched them and those whose labor they profited from. That seems like a screwy but ultimately more than fair deal for the top 10%. Especially when the alternative is pitchforks and torches.

Before you go to the capitalist in their home and torch it or beat them in front of their family, try striking for a while. Usually they soften rather quickly when labour is collectively withheld.

Hence why corporate leadership is salivating over AI.

Re: Google Books (or similar) all book scans – $200k bounty (2025)

#307

Earlier quoted context omitted.

I disagree with your definition of meaningful here. Society's willingness to pay is certainly a signal for meaning in an output but it seems quite inaccurate. Think of the number of artists and thinkers that weren't recognised in their lifetimes, their work was still meaningful but society hadn't discovered it yet. Similarly, there are a number of things that would be incredibly meaningful to all of society (eradicat…

BI does not stop people doing meaningful things. Society will (mostly) reward things which add value. We have a very efficient system for that, and it doesn't go away under BI. We are already spending massive amounts of money on disease, fusion and so on. There's no issue there, and BI doesn't move that needle. At the moment society (especially in the US) operates on a "add value or starve" basis. (That's an over sim…

"add value or starve"

I think this is wrong. It's not about value, it's about being submissive.

Re: Google Books (or similar) all book scans – $200k bounty (2025)

#308
post #98

I think this would cross the line from civil copyright claims into criminal activity https://chatgpt.com/share/6a4970e8-7fe8-83e9-8f81-3aefd76b6b... On another note, if Google's cybersecurity were always one rogue employee away from a massive leak, then it wouldn't be Google. What was the last Google leak you remember, defense in depth people.

AA is an openly criminal organisation. Their attitude to prosecution is "you'll never catch us lol"

Indeed, took a closer look, and the court order that took down their .org domain included "Computer fraud and abuse" claims.

Surprisingly the order is very specific about DNS registrars and authoritative domains not advertising the AA servers, so sharing IP addresses or alternative domain names through other means, like WikiPedia, is not against the order. Which means that nowadays the Wikipedia page works as a pseudo authoritative DNS.

Re: Google Books (or similar) all book scans – $200k bounty (2025)

#309
post #85

Earlier quoted context omitted.

AA compiles from everywhere; LibGen and Z-Lib served as the major sources of books. This has unfortunately led to search results for a particular book containing multiple versions of that book, and it is not readily clear which one is the highest-quality version. A real library would have librarian staff who carefully curate everything, but in the pirate world this isn’t realistic so it just gets all thrown together.…

Surely this is realistic now (or soon) in the form of LLM curation. A few auto-librarians reading everything, looking at different versions side by side, making choices, etc.

LOL, not realistic at all. The differences between book versions lie in more than the raw text. Moreover, an archival project would be loathe to favour or disfavour a version unless an actual human made the call.

Re: Google Books (or similar) all book scans – $200k bounty (2025)

#310
post #41

Earlier quoted context omitted.

https://send.djazz.se/ This is key for getting epubs to your Kobo.

I don't understand what this is doing. Can't you sideload any ebook onto a kobo anyway? Never had an issue on my Clara

Sideload without cables is challenging.

I download epubs from zlib but then they're on my phone and transferring them to my Kobo is arduous. This makes it easy.

Post reply on HN