Live data from Hacker News

Google Books (or similar) all book scans – $200k bounty (2025)

software.annas-archive.gl

31–40 of 368 posts

Re: Google Books (or similar) all book scans – $200k bounty (2025)

#31
post #18

The US should just find a way to quietly share literature access with the Russians, rather than letting piracy be promoted and facilitated for US consumers as freedom-fighter "archiving". Between all the piracy, and all the AI training and the purchase/visitor-circumventing AI services, the practice of writing and publishing genuinely good work is being wiped out. We're killing the goose that lays the eggs, for selfi…

This ship has sailed for academic publications, and academics define that term very liberally because we want to read everything, fiction included. The shadow libraries started off as a way for scholars in ex-Soviet countries in particular (but also India, SE Asia, etc.) to access literature that simply wasn’t available in their country. But the shadow libraries proved so successful and convenient that academics in all countries are using them now, even if they have access to official subscription services. I use AA several times a day and so do the researchers around me in my office; at conferences, if the presenter mentions an interesting publication, the whole room immediately opens AA on their laptops, etc.

Even if projects like AA didn’t have nation-level support, academics would find a way to keep as much of it as possible going. After all, we’re the ones who compiled the bulk of pre-2020 material, and we’re the ones who do all the hard work of scanning from our institutional libraries stuff that doesn’t exist anywhere in digital form.

Re: Google Books (or similar) all book scans – $200k bounty (2025)

#32

Anyone afraid of being laid off at google right now? Perhaps this is a backup :)

I think if you get caught exfiltrating data they'll sue you for much more than $200K.

If your money is in private crypto or offshore you have nothing to worry about.

Re: Google Books (or similar) all book scans – $200k bounty (2025)

#34

Earlier quoted context omitted.

I think if you get caught exfiltrating data they'll sue you for much more than $200K.

If your money is in private crypto or offshore you have nothing to worry about.

Except perhaps jail time.

Lying about your assets to avoid paying a lawful fine is criminal. Just because they can’t see your money doesn’t mean they can’t prove that you have it, and can’t jail you for hiding it to get out paying a fine.

Re: Google Books (or similar) all book scans – $200k bounty (2025)

#35
https://SourceLibrary.org has about 16,000 rare books translated — most for the first time. 50,000 books archived (will be translated when we have $$ for it). More tokens than English Wikipedia and about .75 petabytes.

Not sure if we will qualify for a bounty, but happy to share! Btw, we are looking for funding from small or large donors who want to help us translate the Renaissance…

Re: Google Books (or similar) all book scans – $200k bounty (2025)

#36

Anyone afraid of being laid off at google right now? Perhaps this is a backup :)

I think if you get caught exfiltrating data they'll sue you for much more than $200K.

Copy data into extra large capacity micro sdcard and hide it in your rubiks cube, nobody will suspect a thing

Re: Google Books (or similar) all book scans – $200k bounty (2025)

#37

Curious as to how you would approach this. I have no experience in this area, anyone on this forum willing to share their expertise?

If it works as AA seems to theorize, you'd need to:

  (a) work out how Google books exposes fragments of books, and see if there's a systematic way of using this to get whole books.  For example, a naive approach might be to find any fragment of the book by searching some exact phrase.  Then, you can search for an exact phrase from the start or end of the fragment it gave you, hoping it will show you the previous or next part of the book.  You can then just loop that to get the whole book.

  (b) once you have (a), you need a way of bypassing Google's bot detection/rate limiting.  I don't know what current state of the art is, but there may be a solution for sale out there.  E.g. you pay to receive a cookie or browser state, and use that to fetch the URLs from (a).  Or if you're good/already in the scene, you could do this part yourself.

Re: Google Books (or similar) all book scans – $200k bounty (2025)

#38
post #28

Earlier quoted context omitted.

Not worried about that, you will only have to wait 3-6 months and get a Chinese model just as good.

That’s misunderstanding why these models are behind. A large part of why they’re behind is they aren’t able to do the reinforcement learning post-training steps that takes a pre-trained model and turns it into a frontier model like GPT 5 or Opus. Instead they do their best to recreate these models using distillation. Fundamentally, you can never distill your way to being the teacher, so these approaches will not adva…

That’s not remotely true. They did distillation as a cheap solution to the cold start problem. You need data/trajectories to hill climb to higher capabilities. All large Chinese labs do RLAIF.

Re: Google Books (or similar) all book scans – $200k bounty (2025)

#39
post #28

Earlier quoted context omitted.

Not worried about that, you will only have to wait 3-6 months and get a Chinese model just as good.

That’s misunderstanding why these models are behind. A large part of why they’re behind is they aren’t able to do the reinforcement learning post-training steps that takes a pre-trained model and turns it into a frontier model like GPT 5 or Opus. Instead they do their best to recreate these models using distillation. Fundamentally, you can never distill your way to being the teacher, so these approaches will not adva…

>"they aren’t able to do the reinforcement learning post-training steps"

Not yet.

If there is a need someone will come and fulfill. Personally for me now I do not even want to use top models. Professionally I use AI to help with the coding using Junie agent that comes with IDEs from JetBrains. Junie is told to use Gemini Flash and works fine for what I ("I" being an emphasis here) ask it to do. I tried more advanced models and different vendors only to discover credits going down the toilet without any extra benefit.

Re: Google Books (or similar) all book scans – $200k bounty (2025)

#40
post #14
post #4

Piracy / copyright predictions? The current situation feels untenable with renting. So many regular people I know have learned about VPN, NAS, etc.

Hopefully the guillotines. Look up how much the authors and artists who create the actual work get paid.

Quite a few textbook authors I know are paid well to be part of the whole scheme (kickbacks, forced yearly repurchase for the 'online' component of books, etc). So I think it varies a lot.
Post reply on HN