Google Books (or similar) all book scans – $200k bounty (2025)
91–100 of 368 posts
Re: Google Books (or similar) all book scans – $200k bounty (2025)
#92I live in a country where the selection of available books, especially in English, is very limited. Buying online from foreign markets comes with a long list of administrative hurdles and limits. If it were not for Anna's Archive and Z-Library, I would've never been able to read the books that shaped who I am today, or keep my passion for learning alive. Thanks, AA and ZLib! (Also, thank you to the authors whose book…
Look, fair enough from your perspective. But a lot of those books probably wouldn't exist if the author couldn't make some money from their work. I can't find the post but years ago on Reddit an author posted stats showing when her book turned up pirates online, real sales for it collapsed. Because of this I make a point of buying books, programming books especially. Yes I download pdfs, I use them as previews. This…
I think that's at least a bit debatable. People thought that about (normal) libraries back in the day, but it ended up having the opposite effect.
Not to mention out of print books or academic books which is a big usage of sites like these, since lots of people prefer physical books and only reach for pdfs as a last resort.
Re: Google Books (or similar) all book scans – $200k bounty (2025)
#93It seems like there are some deep pockets funding them.
Re: Google Books (or similar) all book scans – $200k bounty (2025)
#94Re: Google Books (or similar) all book scans – $200k bounty (2025)
#95Re: Google Books (or similar) all book scans – $200k bounty (2025)
#96Earlier quoted context omitted.
I wish an extra capacity SD card was enough, google books holds (probably) an insane numbers of books
Comments on the source mention dataset sizes ranging between 1.5PB and 200PB
Re: Google Books (or similar) all book scans – $200k bounty (2025)
#97Earlier quoted context omitted.
Curious as to what your budget was to get where you are today? That's a lot of tokens. I presume you are using gemini flash?
All the models used are shown with each page of translation and each book has a whole data provenance treatment. You can add it up!
Re: Google Books (or similar) all book scans – $200k bounty (2025)
#98https://chatgpt.com/share/6a4970e8-7fe8-83e9-8f81-3aefd76b6b...
On another note, if Google's cybersecurity were always one rogue employee away from a massive leak, then it wouldn't be Google. What was the last Google leak you remember, defense in depth people.
Re: Google Books (or similar) all book scans – $200k bounty (2025)
#99Earlier quoted context omitted.
Chinese companies giving away expensive models for free is a symptom of the AI bubble, too. It's not a law of nature that they'll always be able to scrounge up the money for yet another training run.
I think it's a deliberate business strategy of commoditization of their complement. China acts like an entire bloc, not as single companies, and they want to monetize hardware.
ByteDance is going the direct-to-consumer route with their Doubao chatbot (the most popular in China, probably thanks to their social media prowess). iFlyTek seems to be angling for enterprise and government use cases, where they already have an in.
The companies that have released weights have in common that they didn't have a monetization channel lined up and their models weren't good enough to make people pay attention with just API access. (You can see with Qwen Max that the calculus can change towards not releasing weights for better models.)
And who exactly among the investors is having their complement commoditized? When Nvidia releases Nemotron, the story is clear, but it's less obvious for say Z.ai's GLM.
Re: Google Books (or similar) all book scans – $200k bounty (2025)
#100Earlier quoted context omitted.
Curious as to what your budget was to get where you are today? That's a lot of tokens. I presume you are using gemini flash?
All the models used are shown with each page of translation and each book has a whole data provenance treatment. You can add it up!
I have seen Gemini costs change quite a bit when processing very similar books from the same series lately, mainly because thinking tokens have increased about 5x. Has that has happened to you as well?
Edit: for ocr I am using about 15k-25k tokens per page, but I have a complex prompt.