Live data from Hacker News

Google Books (or similar) all book scans – $200k bounty (2025)

software.annas-archive.gl

111–120 of 368 posts

Re: Google Books (or similar) all book scans – $200k bounty (2025)

#111

https://SourceLibrary.org has about 16,000 rare books translated — most for the first time. 50,000 books archived (will be translated when we have $$ for it). More tokens than English Wikipedia and about .75 petabytes. Not sure if we will qualify for a bounty, but happy to share! Btw, we are looking for funding from small or large donors who want to help us translate the Renaissance…

Wow this is amazing!

Re: Google Books (or similar) all book scans – $200k bounty (2025)

#112

One of my hopes is that when the AI bubble bursts, some brave person will sneak out a copy of the last frontier model.

If it's a bubble, why do you care about frontier models?

If we had the dotcom bubble, why are you still on the Internet?

Re: Google Books (or similar) all book scans – $200k bounty (2025)

#113
Anna’s came clutch for me yesterday. I spent a few days trying to find a zip file of a CD that came with an old book from early 2000s on programming. One of those Thomson Publishing slap jobs that I actually enjoyed. I checked used copies all of them said does not come with CD. I tried googling around, nothing. LLMs couldn’t find it. ChatGPT kept saying it is on the archive (no it isn’t you useless piece of shit). Anyway, on a whim I went to AA, lo and behold, zip files for both first and second edition. Godsend.

Re: Google Books (or similar) all book scans – $200k bounty (2025)

#114
post #41

Earlier quoted context omitted.

https://send.djazz.se/ This is key for getting epubs to your Kobo.

Thanks, but I don't use e-readers as they are not available here. I've been using MoonReader for many years now and settled on pretty good parameters that make the reading experience very comfortable on both my phone and my tablet.

Moon reader is amazing. I love mine so much I don't see a point of having a separate book reader.

Re: Google Books (or similar) all book scans – $200k bounty (2025)

#115
post #73

Earlier quoted context omitted.

>the practice of writing and publishing genuinely good work is being wiped out. Most of the best literature in the English language was written before modern IP law was even a thing. There's very little good literature written by authors primarily motivated by money.

How much of that literature was written by wealthy landowners who already had little need for money?

Well, you needed the means to get an education, since most of the poor in those days were illiterate, which is something of an impediment to becoming a successful writer.

I can only think of one writer off hand who wasn’t a wealthy landowner, although it is a particularly notable example; that of William Shakespeare.

Shakespeare wasn’t poor (his parents seem to be of upper middle class standing), he was able to get a basic (but not a university) education and then pursue an acting career (with perhaps a side hustle as a teacher). Whatever the case he certainly wasn’t independently wealthy before he started writing, he needed to earn a living.

He did seem to be in it for the money (and fame) since he wasn’t just a writer he was an actor, theatre owner, and something of a celebrity, and he did make enough money to become a wealthy landowner by the time he died.

Re: Google Books (or similar) all book scans – $200k bounty (2025)

#116

So AA is a front for openai?

No, but they openly make a lot of money from selling their library to AI companies. Fast enterprise access to Anna's Archive starts at $100.000

A lot? I would be kind of interested if there were any known figures. Do companies want to be implicated in AA-cooperation in any capacity?

Re: Google Books (or similar) all book scans – $200k bounty (2025)

#117
post #28

Earlier quoted context omitted.

That’s misunderstanding why these models are behind. A large part of why they’re behind is they aren’t able to do the reinforcement learning post-training steps that takes a pre-trained model and turns it into a frontier model like GPT 5 or Opus. Instead they do their best to recreate these models using distillation. Fundamentally, you can never distill your way to being the teacher, so these approaches will not adva…

> you can never distill your way to being the teacher Are you sure? What if you distill from 10 teachers?

In this case all teachers have also learned from each other.

Re: Google Books (or similar) all book scans – $200k bounty (2025)

#118

Earlier quoted context omitted.

No, but they openly make a lot of money from selling their library to AI companies. Fast enterprise access to Anna's Archive starts at $100.000

A lot? I would be kind of interested if there were any known figures. Do companies want to be implicated in AA-cooperation in any capacity?

They likely use intermediary companies, but NVIDIA might have purchased from them directly, I don't remember the full story.

Re: Google Books (or similar) all book scans – $200k bounty (2025)

#119

Who is behind Annas archive, there is a lot of english speakers involved in the team and forums! Anyway as long as buying isn´t owning no issues here.

If no issue there, then why would you ask who is behind it in a public forum?

Re: Google Books (or similar) all book scans – $200k bounty (2025)

#120
post #92

Earlier quoted context omitted.

> But a lot of those books probably wouldn't exist if the author couldn't make some money from their work. I think that's at least a bit debatable. People thought that about (normal) libraries back in the day, but it ended up having the opposite effect. Not to mention out of print books or academic books which is a big usage of sites like these, since lots of people prefer physical books and only reach for pdfs as a…

Can you imagine if we didn’t have libraries and someone tried to create them today? From publishers to right wingers, they would be painted as communist plots to destroy creativity.

The Internet Archive tried, at great cost and peril, to defend its ability to lend books as an online library due to format shift (physical books get first sale doctrine, ebooks are licensed, you cannot own them), and were told no by the system, so “pirating” it is until copyright changes and becomes more reasonable. Disk is cheap, and the Internet global. Global distributed storage system durability and availability is the path to success until laws change imho.

(Archiving culture alone is not the same as also enabling universal access to the culture and knowledge one is acting as custodian for and serving to global citizens)

The Internet Archive has lost its appeal in Hachette vs. Internet Archive - https://news.ycombinator.com/item?id=41447758 - September 2024 (793 comments)

https://archive.org/details/brewsterkahlelongnowfoundation

Totally unrelated: Dweb camp 2026 is coming up for those interested: https://dwebcamp.org/

(no affiliation with any person or entity mentioned in this comment)

Post reply on HN