Live data from Hacker News

Google Books (or similar) all book scans – $200k bounty (2025)

software.annas-archive.gl

91–100 of 368 posts

Re: Google Books (or similar) all book scans – $200k bounty (2025)

#92
post #86

I live in a country where the selection of available books, especially in English, is very limited. Buying online from foreign markets comes with a long list of administrative hurdles and limits. If it were not for Anna's Archive and Z-Library, I would've never been able to read the books that shaped who I am today, or keep my passion for learning alive. Thanks, AA and ZLib! (Also, thank you to the authors whose book…

Look, fair enough from your perspective. But a lot of those books probably wouldn't exist if the author couldn't make some money from their work. I can't find the post but years ago on Reddit an author posted stats showing when her book turned up pirates online, real sales for it collapsed. Because of this I make a point of buying books, programming books especially. Yes I download pdfs, I use them as previews. This…

> But a lot of those books probably wouldn't exist if the author couldn't make some money from their work.

I think that's at least a bit debatable. People thought that about (normal) libraries back in the day, but it ended up having the opposite effect.

Not to mention out of print books or academic books which is a big usage of sites like these, since lots of people prefer physical books and only reach for pdfs as a last resort.

Re: Google Books (or similar) all book scans – $200k bounty (2025)

#93
How is Anna's Archive funded? I see they have memberships, but it's hard to believe that can fund all these bounties - some going into six figures. Ask any FOSS project about funding by that method.

It seems like there are some deep pockets funding them.

Re: Google Books (or similar) all book scans – $200k bounty (2025)

#95

Earlier quoted context omitted.

I wish an extra capacity SD card was enough, google books holds (probably) an insane numbers of books

Comments on the source mention dataset sizes ranging between 1.5PB and 200PB

my guess would be the 7PB mark

Re: Google Books (or similar) all book scans – $200k bounty (2025)

#96

Earlier quoted context omitted.

I wish an extra capacity SD card was enough, google books holds (probably) an insane numbers of books

Comments on the source mention dataset sizes ranging between 1.5PB and 200PB

For 200PB one would need 25kg worth of 2TB microSD cards... that would be lots of Rubik's cubes =P

Re: Google Books (or similar) all book scans – $200k bounty (2025)

#97
post #70

Earlier quoted context omitted.

Curious as to what your budget was to get where you are today? That's a lot of tokens. I presume you are using gemini flash?

All the models used are shown with each page of translation and each book has a whole data provenance treatment. You can add it up!

How do you handle the more densely written pages in script ? I did a very similar exercise OCRing works from this exact collection, but I stuck with the English books for the first pass.

Re: Google Books (or similar) all book scans – $200k bounty (2025)

#98
I think this would cross the line from civil copyright claims into criminal activity

https://chatgpt.com/share/6a4970e8-7fe8-83e9-8f81-3aefd76b6b...

On another note, if Google's cybersecurity were always one rogue employee away from a massive leak, then it wouldn't be Google. What was the last Google leak you remember, defense in depth people.

Re: Google Books (or similar) all book scans – $200k bounty (2025)

#99
post #13
post #8

Earlier quoted context omitted.

Chinese companies giving away expensive models for free is a symptom of the AI bubble, too. It's not a law of nature that they'll always be able to scrounge up the money for yet another training run.

I think it's a deliberate business strategy of commoditization of their complement. China acts like an entire bloc, not as single companies, and they want to monetize hardware.

If you think Chinese companies always act as a bloc, your mental model needs to get about a billion times more detailed. But in this case just a few details may be enough: There are Chinese AI companies that have released LLMs without publishing the weights.

ByteDance is going the direct-to-consumer route with their Doubao chatbot (the most popular in China, probably thanks to their social media prowess). iFlyTek seems to be angling for enterprise and government use cases, where they already have an in.

The companies that have released weights have in common that they didn't have a monetization channel lined up and their models weren't good enough to make people pay attention with just API access. (You can see with Qwen Max that the calculus can change towards not releasing weights for better models.)

And who exactly among the investors is having their complement commoditized? When Nvidia releases Nemotron, the story is clear, but it's less obvious for say Z.ai's GLM.

Re: Google Books (or similar) all book scans – $200k bounty (2025)

#100
post #70

Earlier quoted context omitted.

Curious as to what your budget was to get where you are today? That's a lot of tokens. I presume you are using gemini flash?

All the models used are shown with each page of translation and each book has a whole data provenance treatment. You can add it up!

I don't see raw token counts, just a list of steps and page counts. For example, what is the rough average token count per page in the ocr and in the translation steps for a Greek book?

I have seen Gemini costs change quite a bit when processing very similar books from the same series lately, mainly because thinking tokens have increased about 5x. Has that has happened to you as well?

Edit: for ocr I am using about 15k-25k tokens per page, but I have a complex prompt.

Post reply on HN