Live data from Hacker News

Project Gutenberg – keeps getting better

gutenberg.org

111–120 of 300 posts

Re: Project Gutenberg – keeps getting better

#113
post #95

Earlier quoted context omitted.

that's cool! one of my "pet-ideas" is actually to make an AI-agent that does all that typographical work for any PG book to make it nicely printable without any manual labor whatsoever. Maybe that's doable now ...

That is doable. Most of my work was regexp and repetitive stuff. And the typograhpy stuff is achievable with the current state of the art models. Not that I remember what I did, it was 30 years ago.

Interesting!

Re: Project Gutenberg – keeps getting better

#114

I'm surprised no eBook Reader vendor has a Project Gutenberg "Store." Where you can just browse Gutenberg, find a book, and just grab it down to the reader. Instead, they either are actively hostile (Kindle), or require the use of Calibre (which itself is good, it is just the friction).

Used to be one could sort of get that with the Project Librivox: https://librivox.org/ e-book app Gutebooks (in addition to their audio app), but it seems to have been deprecated (I'm no longer able to connect to the server on my copy (which I only got 'cause there was an in-app purchase to fund Project Librivox). FWIW, Barnes & Noble has been plundering the public domain using a book composition/keying house in the…

>Barnes & Noble has been plundering the public domain using a book composition/keying house in the Philippines to make their public domain books which they make available in their stores

Why is it 'plundering' for B&N to print physical books, transport them to their brick-and-mortar stores to sell? There are real costs associated to doing so. It would not have zero cost for me to print and bind a copy myself at home.

Re: Project Gutenberg – keeps getting better

#115
post #101

Please give me some book recommendations :)

Not a recommendation per se but I used to use Amphetype on Gutenberg texts to practise touch-typing. There's something about writing out a book that hits differently to reading it. You skip less, odd parts stick with you. I think the last one I tried was The Island of Dr Moreau.

Ulnar Nerve Entrapement :/

Re: Project Gutenberg – keeps getting better

#117

I'm surprised no eBook Reader vendor has a Project Gutenberg "Store." Where you can just browse Gutenberg, find a book, and just grab it down to the reader. Instead, they either are actively hostile (Kindle), or require the use of Calibre (which itself is good, it is just the friction).

I've used https://standardebooks.org/ to pull nicely formatted Project Gutenberg books on any e-reader that supports a browser (in my case, Boox). Technically, I can also just directly pull the epub from Project Gutenberg, but sometimes the formatting leaves a lot to be desired. Once you get an e-reader that runs a semi-capable OS (ex - stock android, even an older version), it's hard to go back to something like a k…

To be precise, the vast majority of SE is from Gutenberg, but we also source from Faded Page, Gutenberg Australia, Wikisource and occasionally do our own transcriptions.

Re: Project Gutenberg – keeps getting better

#118
post #2

Hi! I'm one of the programmers at Gutenberg. We've been improving the site a lot over the past few months (and more is coming!). If you haven't visited the page recently, it's worth checking out again: https://www.gutenberg.org/

I can't say for project Gutenberg specifically, but in general a huge issue I see is OCR errors. What do you all do to address OCR?

Check out Distributed Proofreaders: https://pgdp.net

Re: Project Gutenberg – keeps getting better

#120

While PG has probably gotten a lot of use and growth with the growth/maintreaming of the Internet since the 1990s, (TIL) it started back in 1971: > Michael S. Hart began Project Gutenberg in 1971 with the digitization of the United States Declaration of Independence.[5] Hart, a student at the University of Illinois, obtained access to a Xerox Sigma V mainframe computer in the university's Materials Research Lab. […]…

"Project Gutenberg began in 1971 when Michael Hart was given an operator’s account with $100,000,000 of computer time in it by the operators of the Xerox Sigma V mainframe at the Materials Research Lab at the University of Illinois."

https://www.gutenberg.org/about/background/history_and_philo...

Post reply on HN