Live data from Hacker News

Google Books (or similar) all book scans – $200k bounty (2025)

software.annas-archive.gl

61–70 of 368 posts

Re: Google Books (or similar) all book scans – $200k bounty (2025)

#61
post #43

Earlier quoted context omitted.

I think if you get caught exfiltrating data they'll sue you for much more than $200K.

I don't think anybody would do it purely for money. I would rather see someone who is terminally ill and decides to do some "good".

There are not too many mentally-sharp, fully-employed, terminally-ill people that I have met. Even fewer at tech companies.

And even fewer who are single and childless. (Google would likely go after the estate of anyone who did this.)

Re: Google Books (or similar) all book scans – $200k bounty (2025)

#62
post #18

The US should just find a way to quietly share literature access with the Russians, rather than letting piracy be promoted and facilitated for US consumers as freedom-fighter "archiving". Between all the piracy, and all the AI training and the purchase/visitor-circumventing AI services, the practice of writing and publishing genuinely good work is being wiped out. We're killing the goose that lays the eggs, for selfi…

>the practice of writing and publishing genuinely good work is being wiped out.

Most of the best literature in the English language was written before modern IP law was even a thing. There's very little good literature written by authors primarily motivated by money.

Re: Google Books (or similar) all book scans – $200k bounty (2025)

#63
post #41

I live in a country where the selection of available books, especially in English, is very limited. Buying online from foreign markets comes with a long list of administrative hurdles and limits. If it were not for Anna's Archive and Z-Library, I would've never been able to read the books that shaped who I am today, or keep my passion for learning alive. Thanks, AA and ZLib! (Also, thank you to the authors whose book…

https://send.djazz.se/ This is key for getting epubs to your Kobo.

I don't understand what this is doing. Can't you sideload any ebook onto a kobo anyway? Never had an issue on my Clara

Re: Google Books (or similar) all book scans – $200k bounty (2025)

#64
post #8

Earlier quoted context omitted.

Not worried about that, you will only have to wait 3-6 months and get a Chinese model just as good.

Chinese companies giving away expensive models for free is a symptom of the AI bubble, too. It's not a law of nature that they'll always be able to scrounge up the money for yet another training run.

As long as it is in the CCP's national interest to have a frontier model, Chinese companies will have the resources for another training run.

Re: Google Books (or similar) all book scans – $200k bounty (2025)

#65

https://SourceLibrary.org has about 16,000 rare books translated — most for the first time. 50,000 books archived (will be translated when we have $$ for it). More tokens than English Wikipedia and about .75 petabytes. Not sure if we will qualify for a bounty, but happy to share! Btw, we are looking for funding from small or large donors who want to help us translate the Renaissance…

Hey, this looks fascinating!

I can't quickly tell what all you have archived^, but I have some friends who are academic historians who might be interested in certain categories of work (and could help verify some esoteric languages) - is it possible to search by region or language?

Have you reached out to any types of historians WRT the project? It seems like some PhD students might be able to find some projects in this work etc

^ when I looked at the timeline https://sourcelibrary.org/timeline, I got an error

Re: Google Books (or similar) all book scans – $200k bounty (2025)

#66
post #45

Earlier quoted context omitted.

That’s not remotely true. They did distillation as a cheap solution to the cold start problem. You need data/trajectories to hill climb to higher capabilities. All large Chinese labs do RLAIF.

Oh yes, not remotely true. Which is why the frontier labs all have invested heavily in trying to identify and thwart distillers, using known company names / domains to drive their exclusion lists. /s

It's cheaper to distill than to do reinforcement learning, so of course they prefer that, but if it wasn't an option they could just pay up and spend more GPU time on RL.

Re: Google Books (or similar) all book scans – $200k bounty (2025)

#68
post #18

The US should just find a way to quietly share literature access with the Russians, rather than letting piracy be promoted and facilitated for US consumers as freedom-fighter "archiving". Between all the piracy, and all the AI training and the purchase/visitor-circumventing AI services, the practice of writing and publishing genuinely good work is being wiped out. We're killing the goose that lays the eggs, for selfi…

>the practice of writing and publishing genuinely good work is being wiped out. Most of the best literature in the English language was written before modern IP law was even a thing. There's very little good literature written by authors primarily motivated by money.

That's just cultural elitism. I hope you meet someone in your life who finds absolute joy in reading young adult romance novels or D&D fantasy books so you can understand how irrelevant "good" literature is. I love Dostoevsky and Verne (and D&D novels, especially those written by R.A. Salvatore), but I would never judge the modern "IPs" that got my daughter into reading.

> best literature

What does that even mean?

Re: Google Books (or similar) all book scans – $200k bounty (2025)

#69
post #18

The US should just find a way to quietly share literature access with the Russians, rather than letting piracy be promoted and facilitated for US consumers as freedom-fighter "archiving". Between all the piracy, and all the AI training and the purchase/visitor-circumventing AI services, the practice of writing and publishing genuinely good work is being wiped out. We're killing the goose that lays the eggs, for selfi…

>We're killing the goose that lays the eggs, for selfish gain We already did that when the internet collectively agreed decades ago that everything digital should be free for anyone. We're now 20 years downstream of ad-blocking being a virtuous good, and piracy being the ultimate show of liberty, and now suddenly everyone cares about the creator's revenue stream. The mask slipped and unsurprisingly the internet is a…

> Yes, I am talking to you with the 4TB of pirated content, proud of not loading any ads in the last 15 years, and getting enraged over LLM training.

That's oddly-specific :-)

In any case, I have no pirated content that I know off, neither proud nor ashamed of blocking ads[1], but I still get annoyed that a bunch of VCs can use their invested-into companies to launder all the worlds IP, then sell it back to them.

[1] Who feels proud of blocking ads? It's like feeling proud of tying your shoelaces: "Good job, well done, but that's the expectation, son".

Re: Google Books (or similar) all book scans – $200k bounty (2025)

#70

https://SourceLibrary.org has about 16,000 rare books translated — most for the first time. 50,000 books archived (will be translated when we have $$ for it). More tokens than English Wikipedia and about .75 petabytes. Not sure if we will qualify for a bounty, but happy to share! Btw, we are looking for funding from small or large donors who want to help us translate the Renaissance…

Curious as to what your budget was to get where you are today? That's a lot of tokens. I presume you are using gemini flash?
Post reply on HN