Live data from Hacker News

Project Gutenberg – keeps getting better

gutenberg.org

261–270 of 300 posts

Re: Project Gutenberg – keeps getting better

#261
post #259
post #163

Earlier quoted context omitted.

It’s effective against teenagers maybe. Not so much against Amazon, Meta or wherever botnet/crawler is coming out of China these days from up-and-coming AI companies.

Then block all of Amazon, Meta, or wherever botnet/crawling traffic is coming from that doesn't honor robots.txt, sends DDoS reflection traffic, submits SMTP messages (in large volumes, not just probing) for domains they're not authorized for with SPF, or whatever else applies to the protocol you're using If they can't keep their ranges clean to a reasonable degree, their customers will need to move if they want to a…

See the other comments in this thread. The perpetrators are unknown and are jumping between residential IPs. Possibly botnets?

Re: Project Gutenberg – keeps getting better

#262
post #261
post #259

Earlier quoted context omitted.

Then block all of Amazon, Meta, or wherever botnet/crawling traffic is coming from that doesn't honor robots.txt, sends DDoS reflection traffic, submits SMTP messages (in large volumes, not just probing) for domains they're not authorized for with SPF, or whatever else applies to the protocol you're using If they can't keep their ranges clean to a reasonable degree, their customers will need to move if they want to a…

See the other comments in this thread. The perpetrators are unknown and are jumping between residential IPs. Possibly botnets?

Then see my other replies in the thread where I've specifically addressed residential IPs, e.g.: https://news.ycombinator.com/item?id=48163060

Re: Project Gutenberg – keeps getting better

#263
post #258

Earlier quoted context omitted.

Worse than that - even if they would take action, you can't possibly orchestrate filing all of the complaints. It's a drown-in-quicksand problem, you can't fight quicksand one grain at a time.

> you can't possibly orchestrate filing all of the complaints To the ISPs? Each IP range has an abuse email address registered and this is specifically exempt from rate limiting at RIPE's WHOIS server. Not sure how it is in other RIRs but I just happen to know of this policy You can automate the whole thing, provided that you have a reliable way of identifying the undesired traffic which you need anyway for being abl…

See what I wrote above (and let me say I am talking about Project Gutenberg and Distributed Proofreaders here, I am one of the admins on both). A large amount of the hassle traffic we've seen is as I wrote above, the IPs come from everywhere and in many cases, each IP makes a single request and doesn't come back. They change user-agent dynamically, etc, to masquerade as regular traffic. They come from residential, cloud/hyperscale, corporate, educational, government, all the networks, on every continent. This is many thousands of "open a ticket with someone" events per hour territory. It's as difficult to fight as DDoS itself for the same reasons (presumably the harvesting parties know that and that's exactly why this approach is used).

Others online have been writing about their own experience with the same stuff; it's not unique to PG at all, it's everywhere. Talk to anyone that runs a web server and they'll have these stories...

Re: Project Gutenberg – keeps getting better

#264
post #136

Earlier quoted context omitted.

we are having occasional lows in page speed performance due to LARGE amounts of bot traffic. full disclosure - we've not really been able to resolve this fully/well. Let us know if you have a good idea for how to deal with it

I would love it if you could detect AI scraper bots, and feed them AI generated bs instead of the real books...

Cloudflare sells that as a product, they call it Labyrinth IIRC.

Re: Project Gutenberg – keeps getting better

#265
post #177
post #136

Earlier quoted context omitted.

we are having occasional lows in page speed performance due to LARGE amounts of bot traffic. full disclosure - we've not really been able to resolve this fully/well. Let us know if you have a good idea for how to deal with it

If it's purely bot traffic, then Anubis could help You could have seen it on some websites already https://anubis.techaro.lol/

Just to add to the two negative replies, I find Anubis to be the only system that doesn't ever get in the way. My browsers have Javascript enabled and, so far, it never took more than a fraction of a second to complete the checks

Every other system I've run into has constant false positives, e.g. Google captchas will sometimes say I've failed and make me do the hardest level (if it wasn't giving me that already), Cloudflare regularly thinks I'm a bot, Codeberg blocked me before, Github signup captchas used to take ~15 minutes to complete and then still said "well you failed, try again", Github's general rate limiting has false positives (some days I browse a lot, other days little, and on the little days it'll sometimes go "slow down" with no recourse whatsoever, you're just blocked for an indeterminate amount of time), OpenStreetMap blocks my browser at work because I'm using Firefox ESR instead of latest stable and it finds that user agent string to be implausible, whatever the german railway operator uses since a few days is triggering on me constantly, etc.,

etc.,

etc. Constant blocks everywhere.

With Anubis, my understanding is that you do the proof of work (with whatever implementation you like, it doesn't have to be the Javascript one that they provide) and you can move on without ever doing any task yourself. The power consumption is a shame, but so long as attackers aren't even doing this much, the couple Joules it takes doesn't seem to be an issue

Of course, the attackers will evolve, but for now...

Re: Project Gutenberg – keeps getting better

#266
post #2

Hi! I'm one of the programmers at Gutenberg. We've been improving the site a lot over the past few months (and more is coming!). If you haven't visited the page recently, it's worth checking out again: https://www.gutenberg.org/

There should be more books at Gutenberg. Also by the way I just searched for 3d printing and found nothing. Either there are no books, or the search query makes things too complicated, IMO.

As another commenter said PG is almost all books from 95+ years in the past due to copyright law in the US. We partner with a sister organization, the World Library Foundation, who have a self-publishing portal for modern works by authors who wish to put their own work in the public domain. You might want to look there for more modern material. https://self.gutenberg.org

Re: Project Gutenberg – keeps getting better

#267
post #258

Earlier quoted context omitted.

> you can't possibly orchestrate filing all of the complaints To the ISPs? Each IP range has an abuse email address registered and this is specifically exempt from rate limiting at RIPE's WHOIS server. Not sure how it is in other RIRs but I just happen to know of this policy You can automate the whole thing, provided that you have a reliable way of identifying the undesired traffic which you need anyway for being abl…

See what I wrote above (and let me say I am talking about Project Gutenberg and Distributed Proofreaders here, I am one of the admins on both). A large amount of the hassle traffic we've seen is as I wrote above, the IPs come from everywhere and in many cases, each IP makes a single request and doesn't come back. They change user-agent dynamically, etc, to masquerade as regular traffic. They come from residential, cl…

I'm aware, I also host various websites that see an IP do a single request to the most unlikely of deep pages. Usually not hard to correlate with similar surprising requests from the same ISP, though, and that's exactly why it would be useful to talk to them: they know who used that IP address at the given timestamp. If they get a hundred complaints from different websites, the ISP is in the unique position to correlate that and find the subscriber(s) that are problematic

You also don't have to send out 1k support requests per hour. Could trial it with some hosting provider that you expect is responsive and see how it works out

edit: like, I just don't see another solution short of banning being anonymous online. Each site would have to know who you are. Someone has to be able to track it back to a person that is doing the abuse or there can't be any rules that we can apply. Imo it's better if that's the ISP (or VPN provider, say) who already has this information anyway

Re: Project Gutenberg – keeps getting better

#268
post #62

Nice to see so much appreciation for what we do. (I'm the new-ish executive director.) Any wikipedians reading this, the article about PG is... aging. Last I looked, it said we offered Plucker files. @Jseiko has done some nice work.

FYI, I took Plucker out of the lead in November, after a PG volunteer recommended that update on the article talk page. Plucker is currently only mentioned in a sentence about formats offered in 2009.

Happy to make other updates! Writing specific notes on the talk page is helpful.

Re: Project Gutenberg – keeps getting better

#270

Project Gutenberg had (has?) a tendency toward plaintext that always put me off. (And it has been over a decade I'm sure since I explored the site—so I am no doubt now misinformed.) I like a styled formatted book—would prefer PDFs. (I know, not a popular format apparently.) I like the idea of Project Gutenberg but guess I found book scans on archive.org my preference. My go-to example is Lewis Carroll's "Through the…

This is covered in the FAQ - https://www.gutenberg.org/help/faq.html#why-is-project-guten...

And as another person noted, the vast majority of books have HTML, EPUB, Mobi formats. We are also looking at both KEPUB (Kobo) and PDF which will probably come in the future.

Post reply on HN