Earlier quoted context omitted.
It’s effective against teenagers maybe. Not so much against Amazon, Meta or wherever botnet/crawler is coming out of China these days from up-and-coming AI companies.
Then block all of Amazon, Meta, or wherever botnet/crawling traffic is coming from that doesn't honor robots.txt, sends DDoS reflection traffic, submits SMTP messages (in large volumes, not just probing) for domains they're not authorized for with SPF, or whatever else applies to the protocol you're using If they can't keep their ranges clean to a reasonable degree, their customers will need to move if they want to a…
Project Gutenberg – keeps getting better
261–270 of 300 posts
Re: Project Gutenberg – keeps getting better
#262Earlier quoted context omitted.
Then block all of Amazon, Meta, or wherever botnet/crawling traffic is coming from that doesn't honor robots.txt, sends DDoS reflection traffic, submits SMTP messages (in large volumes, not just probing) for domains they're not authorized for with SPF, or whatever else applies to the protocol you're using If they can't keep their ranges clean to a reasonable degree, their customers will need to move if they want to a…
See the other comments in this thread. The perpetrators are unknown and are jumping between residential IPs. Possibly botnets?
Re: Project Gutenberg – keeps getting better
#263Earlier quoted context omitted.
Worse than that - even if they would take action, you can't possibly orchestrate filing all of the complaints. It's a drown-in-quicksand problem, you can't fight quicksand one grain at a time.
> you can't possibly orchestrate filing all of the complaints To the ISPs? Each IP range has an abuse email address registered and this is specifically exempt from rate limiting at RIPE's WHOIS server. Not sure how it is in other RIRs but I just happen to know of this policy You can automate the whole thing, provided that you have a reliable way of identifying the undesired traffic which you need anyway for being abl…
Others online have been writing about their own experience with the same stuff; it's not unique to PG at all, it's everywhere. Talk to anyone that runs a web server and they'll have these stories...
Re: Project Gutenberg – keeps getting better
#264Earlier quoted context omitted.
we are having occasional lows in page speed performance due to LARGE amounts of bot traffic. full disclosure - we've not really been able to resolve this fully/well. Let us know if you have a good idea for how to deal with it
I would love it if you could detect AI scraper bots, and feed them AI generated bs instead of the real books...
Re: Project Gutenberg – keeps getting better
#265Earlier quoted context omitted.
we are having occasional lows in page speed performance due to LARGE amounts of bot traffic. full disclosure - we've not really been able to resolve this fully/well. Let us know if you have a good idea for how to deal with it
If it's purely bot traffic, then Anubis could help You could have seen it on some websites already https://anubis.techaro.lol/
Every other system I've run into has constant false positives, e.g. Google captchas will sometimes say I've failed and make me do the hardest level (if it wasn't giving me that already), Cloudflare regularly thinks I'm a bot, Codeberg blocked me before, Github signup captchas used to take ~15 minutes to complete and then still said "well you failed, try again", Github's general rate limiting has false positives (some days I browse a lot, other days little, and on the little days it'll sometimes go "slow down" with no recourse whatsoever, you're just blocked for an indeterminate amount of time), OpenStreetMap blocks my browser at work because I'm using Firefox ESR instead of latest stable and it finds that user agent string to be implausible, whatever the german railway operator uses since a few days is triggering on me constantly, etc.,
etc.,
etc. Constant blocks everywhere.
With Anubis, my understanding is that you do the proof of work (with whatever implementation you like, it doesn't have to be the Javascript one that they provide) and you can move on without ever doing any task yourself. The power consumption is a shame, but so long as attackers aren't even doing this much, the couple Joules it takes doesn't seem to be an issue
Of course, the attackers will evolve, but for now...
Re: Project Gutenberg – keeps getting better
#266Hi! I'm one of the programmers at Gutenberg. We've been improving the site a lot over the past few months (and more is coming!). If you haven't visited the page recently, it's worth checking out again: https://www.gutenberg.org/
There should be more books at Gutenberg. Also by the way I just searched for 3d printing and found nothing. Either there are no books, or the search query makes things too complicated, IMO.
Re: Project Gutenberg – keeps getting better
#267Earlier quoted context omitted.
> you can't possibly orchestrate filing all of the complaints To the ISPs? Each IP range has an abuse email address registered and this is specifically exempt from rate limiting at RIPE's WHOIS server. Not sure how it is in other RIRs but I just happen to know of this policy You can automate the whole thing, provided that you have a reliable way of identifying the undesired traffic which you need anyway for being abl…
See what I wrote above (and let me say I am talking about Project Gutenberg and Distributed Proofreaders here, I am one of the admins on both). A large amount of the hassle traffic we've seen is as I wrote above, the IPs come from everywhere and in many cases, each IP makes a single request and doesn't come back. They change user-agent dynamically, etc, to masquerade as regular traffic. They come from residential, cl…
You also don't have to send out 1k support requests per hour. Could trial it with some hosting provider that you expect is responsive and see how it works out
edit: like, I just don't see another solution short of banning being anonymous online. Each site would have to know who you are. Someone has to be able to track it back to a person that is doing the abuse or there can't be any rules that we can apply. Imo it's better if that's the ISP (or VPN provider, say) who already has this information anyway
Re: Project Gutenberg – keeps getting better
#268Nice to see so much appreciation for what we do. (I'm the new-ish executive director.) Any wikipedians reading this, the article about PG is... aging. Last I looked, it said we offered Plucker files. @Jseiko has done some nice work.
Happy to make other updates! Writing specific notes on the talk page is helpful.
Re: Project Gutenberg – keeps getting better
#269Re: Project Gutenberg – keeps getting better
#270Project Gutenberg had (has?) a tendency toward plaintext that always put me off. (And it has been over a decade I'm sure since I explored the site—so I am no doubt now misinformed.) I like a styled formatted book—would prefer PDFs. (I know, not a popular format apparently.) I like the idea of Project Gutenberg but guess I found book scans on archive.org my preference. My go-to example is Lewis Carroll's "Through the…
And as another person noted, the vast majority of books have HTML, EPUB, Mobi formats. We are also looking at both KEPUB (Kobo) and PDF which will probably come in the future.