Live data from Hacker News

AI crawlers need to be more respectful

about.readthedocs.com

61–70 of 128 posts

Re: AI crawlers need to be more respectful

#61
post #56
post #50

HellPot: https://github.com/yunginnanet/HellPot > Clients (hopefully bots) that disregard robots.txt and connect to your instance of HellPot will suffer eternal consequences. HellPot will send an infinite stream of data that is just close enough to being a real website that they might just stick around until their soul is ripped apart and they cease to exist.

This doesn't solve their bandwidth costs which is their real problem with these bots.

My 100 mbps upload bandwidth at home is free (apart from the monthly 35€ payment). Useless bots will get stuck downloading from me instead of hogging readthedocs.

Re: AI crawlers need to be more respectful

#64
post #58
post #35

Not just AI: here is my current side-quest: https://www.earth.org.uk/RSS-efficiency.html Over 99% of the bandwidth (and CPU) taken by the biggest podcast / music services simply on polling feeds is completely unnecessary. But ofc pointing this out to them gets some sort of "oh this is normal, we don't care" response because they are big enough to know that eg podcasters need them.

Can you insert a podcast item that says "you are spamming our server - please stop"?

Not easily on a static server and not without the risk of annoying an actual real human listener!

If you look at the "Defences" section: https://www.earth.org.uk/RSS-efficiency.html#Hints you'll see there are some things that can be done, such as randomly rejecting a large fraction of requests that don't allow compression (gzip is madly effective on many feed files: it's rude for a client not to allow it). But all these measures take effort to set up, and don't stop the bad bots making the request 100s of times too often. Just responding to each stupid request forces a flurry of packets and wakes up and uses CPU...

Re: AI crawlers need to be more respectful

#65

Earlier quoted context omitted.

Respect to them for not naming names, that's the classy move, but I wouldn't blame them if they did.

Why is it "classy" to not name names when a business (which likely holds it self out there as reputable) behaves badly, especially when it behaves in a way that costs you money? Everyone is so vague and coy. These companies are being abusive and reckless. Name and shame!

In the article, they mention that they are working with the crawling company to be reimbursed for the download costs.

Naming and shaming the company while you're trying to work with them is a real good way to not get what you want.

Re: AI crawlers need to be more respectful

#66

Earlier quoted context omitted.

Respect to them for not naming names, that's the classy move, but I wouldn't blame them if they did.

Why is it "classy" to not name names when a business (which likely holds it self out there as reputable) behaves badly, especially when it behaves in a way that costs you money? Everyone is so vague and coy. These companies are being abusive and reckless. Name and shame!

[deleted]

Re: AI crawlers need to be more respectful

#69
post #10

"One crawler downloaded 73 TB of zipped HTML files in May 2024, with almost 10 TB in a single day. This cost us over $5,000 in bandwidth charges, and we had to block the crawler." Wow. That's some seriously disrespectful crawling.

10TB/day is, roughly, a single saturated 1Gbit link. In technical bandwidth terms, that is the square root of fuck all.

The crazy thing here is that the target site is content to pay ~$700 for the volume of traffic that you can move through a single teeny-tiny included-at-no-extra-charge cat5 Ethernet link in one single day. And apparently, they're going to continue doing so.

Post reply on HN