Live data from Hacker News

Google Books (or similar) all book scans – $200k bounty (2025)

software.annas-archive.gl

331–340 of 368 posts

Re: Google Books (or similar) all book scans – $200k bounty (2025)

#332

Earlier quoted context omitted.

> Knowledge should be free. It was never created in a vacuum. This is a common perspective on HN, but it's so jarring. Someone violates an open-source license and we grab our pitchforks. Someone pirates books and it's fine - really, the authors should be thanking us. Good books are incredibly challenging to write, more so than good software. It's not like you grab Harry Potter and say "I'm just gonna change character…

If this is not bad faith argument, then I don't what is. When someone is violating an OSS licence, they are doing it for commercial gains and monetary profit. Nobody is angry at someone using FOSS software for himself with no money getting involved. As opposed to that, books, movies are pirated for personal consumption. Not monetary gains. If someone bought a $30 book, and then ran a BaaS with millions of VC money in…

The original thought on copyright was that facts should not be copyrightable. This should be extended to all science and all non-fiction. Also the original thought was that parody should not be copyrightable. Most fiction is formula-based, so not too original there, either.

Re: Google Books (or similar) all book scans – $200k bounty (2025)

#333
post #322
post #297

Earlier quoted context omitted.

Yes You could look it up and see for yourself. It's applicable to all sorts of collections such as map data and I'd presume also a book database but IANAL so best if you see for yourself - my intention was just to point out (what I hope is) the applicable legal principle for anyone curious about this

I don't know about that ... HN is a conversation. It's not a very interesting conversation when you say a few words and then refuse to say more, telling the other people at the table to look it up. I could look up lots and lots of things to see for myself. I wouldn't be reading your comment, and there is not enough time in the universe. Readers don't have time to look up everything other people post. It's up to the c…

Fair enough. Not everyone always has time to reiterate Wikipedia on what seems to me a relatively well-known concept, at least among people to whom this is relevant (e.g. me as an OpenStreetMap contributor, database copyright is like the first thing you learn after making your first edit because that's why you can't copy data from just any source). There's always a lucky ten thousand around I suppose :). Maybe I should have at least cited something

Re: Google Books (or similar) all book scans – $200k bounty (2025)

#334
post #7

Some more interesting bounties they offer: https://software.annas-archive.gl/AnnaArchivist/annas-archiv... > Purchase all Library of Congress MARC datasets — $3,000 bounty > English Wikipedia pages about relevant institutions — up to $100 per new page > Internet Archive Digital Lending — $5000 per 1 million pdf files > Text version of our full library — $20,000 ...

Up to 500k for OPSEC failures is interesting, as well. It gives me hope that there are wealthy individuals contributing to sharing books, or many small donations. https://software.annas-archive.gl/AnnaArchivist/annas-archiv...

Anna's Archive likely profited immensely with all the AI labs creating models. They have a dedicated FTP for these companies.

Re: Google Books (or similar) all book scans – $200k bounty (2025)

#335
post #121

Earlier quoted context omitted.

So Anna's Archive is in some ways a front for AI companies, gathering the sources they can't get themselves?

They could get access for free through torrents.

Meta did, numerous Meta IP's appeared in the torrent swarms.

The problem with using the torrents is they are slow, most of the larger sets have <3 seeders, with many dropping in and out, on home or slow connections. I would imagine other companies have learned from Meta's mistake and don't want to appear in the swarm either, which is why direct access is preferred. 100k for unlimited books access is nothing compared to the other costs these labs incur.

Re: Google Books (or similar) all book scans – $200k bounty (2025)

#337

Earlier quoted context omitted.

They could get access for free through torrents.

Meta did, numerous Meta IP's appeared in the torrent swarms. The problem with using the torrents is they are slow, most of the larger sets have <3 seeders, with many dropping in and out, on home or slow connections. I would imagine other companies have learned from Meta's mistake and don't want to appear in the swarm either, which is why direct access is preferred. 100k for unlimited books access is nothing compared…

I'm under the impression that hiding IP addresses is easier than financial transactions, but both should be easy for a trillion dollar company. But I suppose they could use a shady intermediary company with a don't ask don't tell policy which would make high-speed access the final result.

Re: Google Books (or similar) all book scans – $200k bounty (2025)

#338

Earlier quoted context omitted.

Up to 500k for OPSEC failures is interesting, as well. It gives me hope that there are wealthy individuals contributing to sharing books, or many small donations. https://software.annas-archive.gl/AnnaArchivist/annas-archiv...

Anna's Archive likely profited immensely with all the AI labs creating models. They have a dedicated FTP for these companies.

They must have had decent funding to get started, though. Pretty sure they started before the AI boom.

Re: Google Books (or similar) all book scans – $200k bounty (2025)

#339
post #336

Earlier quoted context omitted.

Why would a logical actor pay for something infinitely duplicatable?

[flagged]

We've banned this account for repeatedly breaking the site guidelines and ignoring our requests to stop.

If you don't want to be banned, you're welcome to email hn@ycombinator.com and give us reason to believe that you'll follow the rules in the future. They're here: https://news.ycombinator.com/newsguidelines.html.

Re: Google Books (or similar) all book scans – $200k bounty (2025)

#340

Earlier quoted context omitted.

But then we step back a little further and ask what this thing that is called property, why should any human be granted any beyond what actually constitute them as an entity of their own. What matter at the end of the day is not what the document pretend about who possess what, but how people feel in their life, what they can access to, and what they are bared to access for which actual reason. It can't even be purel…

Property may be a social construct, but the costs of living are not. You can question ownership in the abstract, and I am not even against that conversation. But that does not answer the actual point here. We still live in a world where food, rent, healthcare, clothing, hygiene, servers, tools, and time all cost money. So if someone gives something away out of kindness, access, public benefit, or community spirit, th…

> We still live in a world where food, rent, healthcare, clothing, hygiene, servers, tools, and time all cost money.

We, as "anyone online able to read this", certainly are pushed by social pressure to consider world through that kind of lenses. Not sure money mean anything for any animist tribe isolated in Amazonia.

We can even consider costs, without talking about money. We can for example deem attention span, ecological impact, human relationship, emotional burden. So many things that money tend to weight zero in its process to flatten judgements to dull "well ordered" scalars.

Post reply on HN