Live data from Hacker News

AI crawlers need to be more respectful

about.readthedocs.com

111–120 of 128 posts

Re: AI crawlers need to be more respectful

#112
post #35

Not just AI: here is my current side-quest: https://www.earth.org.uk/RSS-efficiency.html Over 99% of the bandwidth (and CPU) taken by the biggest podcast / music services simply on polling feeds is completely unnecessary. But ofc pointing this out to them gets some sort of "oh this is normal, we don't care" response because they are big enough to know that eg podcasters need them.

I run pinecast.com. If there was a leaderboard for hn users serving XML, I'd almost certainly be in the top five. I don't disagree with your post. But: RSS downloads are at an all time low, and that's a bad thing. They're at an all time low because Spotify and Apple both fetch feeds from centralized servers. 1000 subscribers no longer means 24000ish daily feed fetches, it means 48. With keep alive or H2, these servic…

Do you have a CDN between you and Apple / Spotify? Because if you do I think that Apple/Spotify are polling that CDN every few minutes and the CDN is having its bandwidth wasted invisibly, but presumably priced in.

Also I agree that the re-centralisation is a bad thing, mainly.

(I'd like to move to email to discuss this further, if possible: I have an arXiv paper to write!)

Re: AI crawlers need to be more respectful

#113

Earlier quoted context omitted.

I mean, blocking random IP ranges is your prerogative if you don't want to have customers. The scrapers will find ways around, while actual users will be unable to use your site. Residential proxies are something like $5 for 1000.

True, note domestic ISP IP ranges are published, and unless one deals internationally... don't bother serving people that will never buy anything from your firm anyways. Domestic "Users" functioning as proxies will be tripping usage limits, and getting temporarily banned. Google does this by the way, try hammering their services and find out what happens. Context cookies also immediately flag egregious multi-user rou…

Those of us running sites for public information rather than sales cannot make the simple cut-off that you do.

And again, none of this is simple. It has taken me a few weeks to establish a usage mechanism that does catch the worst feed pullers, but it still can hurt legit new users. That is an opportunity cost.

Re: AI crawlers need to be more respectful

#114

Earlier quoted context omitted.

Naming names isn't really required. The hosts have a $5,000 bandwidth fee, but so do the consumers. There's maybe 10 companies with the financial & compute resources to let a $5,000-per-month-per-website bug run rampant before taking the harvesting service offline. Meta/Google/Whoever may benefit from economies of scale, so they're not seeing the full $5,000 their side, but they're hitting tens of thousands of sites…

You know you can hit the data rate they were complaining about by using a residential fiber connection, right? 10 TB per day is about 1 Gigabit continuous if I’m not mistaken. There are probably millions of people that could to this if they wanted to.

There's millions of people who could do that to an individual website. There are remarkably few organisations who could do that simultaneously across the top 100,000 or so sites on the internet, which is how readthedocs has encountered this issue.

Re: AI crawlers need to be more respectful

#115
> "One crawler downloaded 73 TB of zipped HTML files in May 2024, with almost 10 TB in a single day. This cost us over $5,000 in bandwidth charges, and we had to block the crawler."

Invoice the abusers.

They're rolling in investor hype money, and they're obviously not spending it on competent developers if their bots behave like this, so there should be plenty left to cover costs.

Re: AI crawlers need to be more respectful

#116

Earlier quoted context omitted.

True, note domestic ISP IP ranges are published, and unless one deals internationally... don't bother serving people that will never buy anything from your firm anyways. Domestic "Users" functioning as proxies will be tripping usage limits, and getting temporarily banned. Google does this by the way, try hammering their services and find out what happens. Context cookies also immediately flag egregious multi-user rou…

Those of us running sites for public information rather than sales cannot make the simple cut-off that you do. And again, none of this is simple. It has taken me a few weeks to establish a usage mechanism that does catch the worst feed pullers, but it still can hurt legit new users. That is an opportunity cost.

One must assume most user IP edge proxies are compromised hosts. If someone paid for that list they were almost certainly conned, as the black hats regularly publish that content on their forums. These folks want as many users as possible in order to hide their nuisance traffic origin in the traffic noise.

Allowing users known to have an active RAT or their "proxy friends" on a commercial site is not helping anyone.... especially the victims.

https://www.youtube.com/watch?v=aCbfMkh940Q

Worth studying the problem from time to time when you get bored of the antics.

These folks are generally uninterested in positively contributing to any community, but rather show up to cause trouble for fun and profit.

User API quotas are popular for a reason. =3

Re: AI crawlers need to be more respectful

#117

Earlier quoted context omitted.

I mean, blocking random IP ranges is your prerogative if you don't want to have customers. The scrapers will find ways around, while actual users will be unable to use your site. Residential proxies are something like $5 for 1000.

Exactly, I'm banned or captchaed from half of all web sites these days, because of the AI.

Try updating your web browser, as sites often flag outdated user agent strings hard-coded in many bots/spiders.

Have a nice day, =3

Re: AI crawlers need to be more respectful

#118

Earlier quoted context omitted.

I run pinecast.com. If there was a leaderboard for hn users serving XML, I'd almost certainly be in the top five. I don't disagree with your post. But: RSS downloads are at an all time low, and that's a bad thing. They're at an all time low because Spotify and Apple both fetch feeds from centralized servers. 1000 subscribers no longer means 24000ish daily feed fetches, it means 48. With keep alive or H2, these servic…

Do you have a CDN between you and Apple / Spotify? Because if you do I think that Apple/Spotify are polling that CDN every few minutes and the CDN is having its bandwidth wasted invisibly, but presumably priced in. Also I agree that the re-centralisation is a bad thing, mainly. (I'd like to move to email to discuss this further, if possible: I have an arXiv paper to write!)

The data I'm giving you is based on logs from the CDN. Most feeds are checked by Apple and Spotify every hour, but usually it's less frequently rather than more: shows that haven't been published to in a year or more might see very infrequent feed checks.

Re: AI crawlers need to be more respectful

#119
post #86

Earlier quoted context omitted.

Why are you ending all your messages with =3 ?

https://www.jpl.nasa.gov/images/pia22092-arp-142-the-penguin... Don't worry about it friend =3

Read that before, read it again to make sure I didn't miss anything. The lack of clarity here is disappointing and only asks more questions than it answers.

Re: AI crawlers need to be more respectful

#120

Earlier quoted context omitted.

https://www.jpl.nasa.gov/images/pia22092-arp-142-the-penguin... Don't worry about it friend =3

Read that before, read it again to make sure I didn't miss anything. The lack of clarity here is disappointing and only asks more questions than it answers.

Don't worry about it friend =3
Post reply on HN