Live data from Hacker News

Project Gutenberg – keeps getting better

gutenberg.org

211–220 of 300 posts

Re: Project Gutenberg – keeps getting better

#211
All the books should be there. I understand that current society has restrictions, what with near infinite copyright and other shenanigans - but I don't see any of these as reason to hide information from mankind. Eventually we'll free all the information. Remuneration will have to occur in other ways than the current status quo.

Re: Project Gutenberg – keeps getting better

#212
post #2

Hi! I'm one of the programmers at Gutenberg. We've been improving the site a lot over the past few months (and more is coming!). If you haven't visited the page recently, it's worth checking out again: https://www.gutenberg.org/

There should be more books at Gutenberg.

Also by the way I just searched for 3d printing and found nothing. Either there are no books, or the search query makes things too complicated, IMO.

Re: Project Gutenberg – keeps getting better

#213
post #136

Earlier quoted context omitted.

we are having occasional lows in page speed performance due to LARGE amounts of bot traffic. full disclosure - we've not really been able to resolve this fully/well. Let us know if you have a good idea for how to deal with it

I would love it if you could detect AI scraper bots, and feed them AI generated bs instead of the real books...

This is very, very, very dangerous.

Occasionally, you misclassify a real user as a bot, and then your reputation is ruined forever.

The official Polish train schedules website did this recently, feeding incorrect departure and arrival times to IP addresses known for aggressive scraping, without taking CGNAT into account. People... have noticed[1].

[1] (Polish) https://zaufanatrzeciastrona.pl/post/kto-i-dlaczego-losuje-w...

Re: Project Gutenberg – keeps getting better

#214
post #143
post #136

Earlier quoted context omitted.

we are having occasional lows in page speed performance due to LARGE amounts of bot traffic. full disclosure - we've not really been able to resolve this fully/well. Let us know if you have a good idea for how to deal with it

I'm only a small-scale sysadmin but the way that I understand the internet is that you send abuse notifications to the IP address block owner and, if it doesn't get resolved, you block. The whois/rdap database reveals which IPs all belong to the same hosting provider or ISP, so you can summarize that all to one list of IP addrs + timestamps per some time period The ISP actually knows which subscriber is on that line,…

The problem with this approach is that modern scrapers use hordes of residential proxies and quickly rotate through IP addresses which belong to ASes you get a lot of real traffic from. There's nothing you can do if the ISP won't take any action against the customer.

Re: Project Gutenberg – keeps getting better

#216
post #103

Worth mentioning the Project Gutenberg ZIMs. You can download the entire ENglish Gutenberg corpus for about 60GB (English Wikipedia ZIM complete with images is ~120GB): https://ebookfoundation.org/openzim.html

Like the Project Gutenberg collection on archive.org, the ZIMs are only current up to 2018.

Re: Project Gutenberg – keeps getting better

#218

I'm surprised no eBook Reader vendor has a Project Gutenberg "Store." Where you can just browse Gutenberg, find a book, and just grab it down to the reader. Instead, they either are actively hostile (Kindle), or require the use of Calibre (which itself is good, it is just the friction).

You can download books directly from the Project Gutenberg website using the web browser on most eBook readers - even the Kindle supports it.

Yep! This is how I get all my books on Kindle! For me, I choose the 'older Kindles' option and it downloads directly to my homepage.

Re: Project Gutenberg – keeps getting better

#219
post #2

Hi! I'm one of the programmers at Gutenberg. We've been improving the site a lot over the past few months (and more is coming!). If you haven't visited the page recently, it's worth checking out again: https://www.gutenberg.org/

The biggest lever: make the reading experience great. https://www.gutenberg.org/cache/epub/245/pg245-images.html is still hard to read: lines are tooo long (macbook), no great way for pagination/remembering where I was, notes

Re: Project Gutenberg – keeps getting better

#220

Earlier quoted context omitted.

Found: it's a sentence from 2020, and PG decided not to appeal (!?) Full story (in Italian) at https://www.wired.it/internet/web/2020/06/30/progetto-gutenb...

Seems like a case for HTTP 451 (Unavailable for Legal Reasons) rather than 404.

HTTP 666 (We're evil) seems more fitting here.
Post reply on HN