Live data from Hacker News

Project Gutenberg – keeps getting better

gutenberg.org

241–250 of 300 posts

Re: Project Gutenberg – keeps getting better

#241
post #2

Hi! I'm one of the programmers at Gutenberg. We've been improving the site a lot over the past few months (and more is coming!). If you haven't visited the page recently, it's worth checking out again: https://www.gutenberg.org/

The biggest lever: make the reading experience great. https://www.gutenberg.org/cache/epub/245/pg245-images.html is still hard to read: lines are tooo long (macbook), no great way for pagination/remembering where I was, notes

Lines aren't too long. They look great on all my devices.

Use ⌘ + + until you get the line length you like.

Re: Project Gutenberg – keeps getting better

#242
post #234

Earlier quoted context omitted.

I love, love, looove the fact that I can have a book's html version on project gutenberg bookmarked and continue to read across devices without ever having to login. I use the browser's inbuilt capability extensively to enhance my reading experience (fonts, backgrounds, text to speech, print formatting, share snippets). None of this is a good experience with pdf, epub or any other format. I've read more (meaningful)…

Interesting. Do you "just" use the browser's built-in capabilities or also some browser extensions?

I just use the built-in capabilities these days as everything that I would need is in there. This was not true many years ago when I did use some browser extensions.

Re: Project Gutenberg – keeps getting better

#244

I love Project Gutenberg, don't get me wrong... but frankly, Anna's is better.

I came here to post something similar. PG is perhaps still important as an archive of proofread OCRed public-domain material, but for ordinary people, the shadow libraries have vastly more stuff. After all, readers don’t want their reading to be limited to what was published before a copyright cutoff date many decades ago.

Re: Project Gutenberg – keeps getting better

#245
post #2

Hi! I'm one of the programmers at Gutenberg. We've been improving the site a lot over the past few months (and more is coming!). If you haven't visited the page recently, it's worth checking out again: https://www.gutenberg.org/

There should be more books at Gutenberg. Also by the way I just searched for 3d printing and found nothing. Either there are no books, or the search query makes things too complicated, IMO.

Gutenberg is nearly all books that have lapsed into the US public domain by dint of being published 95+ years in the past. Which broadly explains why you hit nothing for 3d printing.

Re: Project Gutenberg – keeps getting better

#246
post #155

Earlier quoted context omitted.

It looks like the issue was that, in Italy, copyright expires 70 years after the death of the author or the first translator of a work.

PG works based on US copyright law. And as I understand it that's also 70 years after author/translator death. My gut feeling is that if anyone tried hard enough this ban could probably get lifted

That’s only the case for works published after the mid-70s. For works published before (which is all current PD books in the US), it’s 95 years after the date of publication, with a few exceptions where people failed to file renewal notices.

Re: Project Gutenberg – keeps getting better

#247
post #85
post #46

Earlier quoted context omitted.

I like plain text. You can always post process it into any other format you prefer.

it's also very "accessible" - good for assistive technologies and people with "ou-of-the-ordinary" requirements

Well, the problem is that you lose then all the semantic information that was encoded into the HTML or ePub versions. Those tend to be better for assistive tech users.

Re: Project Gutenberg – keeps getting better

#248
post #131
post #2

Hi! I'm one of the programmers at Gutenberg. We've been improving the site a lot over the past few months (and more is coming!). If you haven't visited the page recently, it's worth checking out again: https://www.gutenberg.org/

Huh that's interesting: 4.5 seconds for the TCP handshake and an additional 9.2 seconds for the TLS handshake. Is this some kind of captcha, since most bots would disconnect before that, so if you complete it once then it knows you're good? (Until the bots catch on of course, but so long as it works it's relatively unintrusive and not discriminatory against uncommon client software (that is, non-Chrome/ium).) The res…

traffic yesterday ~20% more than recent average. 4971601 sessions 177 robots 863462 robot files 3390115 user files 20.30% robot files (robots id'd based on requests/ip address) 5 apache servers for static content, 1 CherryPy server for dynamic content hosted at iBiblio.

Re: Project Gutenberg – keeps getting better

#249
post #149
post #136

Earlier quoted context omitted.

we are having occasional lows in page speed performance due to LARGE amounts of bot traffic. full disclosure - we've not really been able to resolve this fully/well. Let us know if you have a good idea for how to deal with it

Do you host a torrent? I have about 50k of the books, I would have used a torrent of just the txt files if it was prominent.

we have a tarball of all text files - link posted somewhere here

Re: Project Gutenberg – keeps getting better

#250
post #161
post #2

Hi! I'm one of the programmers at Gutenberg. We've been improving the site a lot over the past few months (and more is coming!). If you haven't visited the page recently, it's worth checking out again: https://www.gutenberg.org/

As long as you're taking suggestions, since many of the books are quite old, adding a publication date or date range to the search functionality might be nice. I personally would find it very useful since I have a tendency to look for things that are older than year _x_ when researching various things. Thanks for all the effort put into the site!

only 20% of our books have original publication data in the db. We have a project to add another 40% or so from another database, let us know if you want to help.
Post reply on HN