Hi! I'm one of the programmers at Gutenberg. We've been improving the site a lot over the past few months (and more is coming!). If you haven't visited the page recently, it's worth checking out again: https://www.gutenberg.org/
The biggest lever: make the reading experience great. https://www.gutenberg.org/cache/epub/245/pg245-images.html is still hard to read: lines are tooo long (macbook), no great way for pagination/remembering where I was, notes
Project Gutenberg – keeps getting better
231–240 of 300 posts
Re: Project Gutenberg – keeps getting better
#232Looks like the top downloaded book yesterday[0] was Concrete Construction: Methods and Costs by Gillette and Hill.[1] Beat out Moby Dick, Count of Monte Cristo, Frankenstien, Romeo and Juliet, and others. > 23644 downloads in the last 30 days. I wonder if this is bot behavior? 23k downloads feels like a lot? [0] https://www.gutenberg.org/browse/scores/top [1] https://www.gutenberg.org/ebooks/24855
For context, here is the first paragraph of the book's preface: How best to perform construction work and what it will cost for materials, labor, plant and general expenses are matters of vital interest to engineers and contractors. This book is a treatise on the methods and cost of concrete construction. No attempt has been made to present the subject of cement testing which is already covered by Mr. W. Purves Taylo…
Re: Project Gutenberg – keeps getting better
#233Earlier quoted context omitted.
I'm only a small-scale sysadmin but the way that I understand the internet is that you send abuse notifications to the IP address block owner and, if it doesn't get resolved, you block. The whois/rdap database reveals which IPs all belong to the same hosting provider or ISP, so you can summarize that all to one list of IP addrs + timestamps per some time period The ISP actually knows which subscriber is on that line,…
The problem with this approach is that modern scrapers use hordes of residential proxies and quickly rotate through IP addresses which belong to ASes you get a lot of real traffic from. There's nothing you can do if the ISP won't take any action against the customer.
Re: Project Gutenberg – keeps getting better
#234Project Gutenberg had (has?) a tendency toward plaintext that always put me off. (And it has been over a decade I'm sure since I explored the site—so I am no doubt now misinformed.) I like a styled formatted book—would prefer PDFs. (I know, not a popular format apparently.) I like the idea of Project Gutenberg but guess I found book scans on archive.org my preference. My go-to example is Lewis Carroll's "Through the…
I love, love, looove the fact that I can have a book's html version on project gutenberg bookmarked and continue to read across devices without ever having to login. I use the browser's inbuilt capability extensively to enhance my reading experience (fonts, backgrounds, text to speech, print formatting, share snippets). None of this is a good experience with pdf, epub or any other format. I've read more (meaningful)…
Re: Project Gutenberg – keeps getting better
#235All the books should be there. I understand that current society has restrictions, what with near infinite copyright and other shenanigans - but I don't see any of these as reason to hide information from mankind. Eventually we'll free all the information. Remuneration will have to occur in other ways than the current status quo.
Re: Project Gutenberg – keeps getting better
#236Earlier quoted context omitted.
The biggest lever: make the reading experience great. https://www.gutenberg.org/cache/epub/245/pg245-images.html is still hard to read: lines are tooo long (macbook), no great way for pagination/remembering where I was, notes
Firefox's reader mode works amazingly for these situations.
I got it to a prototype level but then shelved it after having difficulty getting good results with various test datasets. Probably would make a fantastic ereader though
Re: Project Gutenberg – keeps getting better
#237Hi! I'm one of the programmers at Gutenberg. We've been improving the site a lot over the past few months (and more is coming!). If you haven't visited the page recently, it's worth checking out again: https://www.gutenberg.org/
Great project. Are many of the books in a format that can easily be converted into audio? Is there a way to search for them, and information on what software your readers find useful for this purpose? (Note: A lot of print media these days has switched to far-to-small font-sizes. Less of a problem for (zoomable) digital media, but for many that's still a barrier.)
human-read: https://www.gutenberg.org/browse/categories/1
computer-generated: https://www.gutenberg.org/browse/categories/2
IIRC many of the human-generated ones come from LibriVox, many of the computer-generated ones came from a collaboration with Microsoft.
Re: Project Gutenberg – keeps getting better
#238Earlier quoted context omitted.
Check out Distributed Proofreaders: https://pgdp.net
I didn't realized DP was still around. I used to do it quite a bit, 15 years ago, but OCR has improved considerably since then.
Re: Project Gutenberg – keeps getting better
#239Earlier quoted context omitted.
Not the GP, but I also have mixed feelings about Standard Ebooks. They modernise texts for American readers. This means changing the punctuation, merging some words, altering the syntax, etc. When I read an old novel, written two centuries ago in England, the little differences to modern English are part of the charm, and I certainly don't want any Americanism mixed in. For one of my favorite novels, The Forsyte saga…
SE editor in chief here. What you describe is incorrect. The only thing we do is very light sound-alike spelling modernization, like "to-night" -> "tonight". We do not do things like change from en-GB to en-US, replace old words with different modern words, or change text for "American readers", whatever that means. I have no idea where you got that impression. I personally worked on the Forsyte saga. If you think so…
Curious. Why even bother?
Re: Project Gutenberg – keeps getting better
#240Good job