AI crawlers need to be more respectful
111–120 of 128 posts
Re: AI crawlers need to be more respectful
#112Not just AI: here is my current side-quest: https://www.earth.org.uk/RSS-efficiency.html Over 99% of the bandwidth (and CPU) taken by the biggest podcast / music services simply on polling feeds is completely unnecessary. But ofc pointing this out to them gets some sort of "oh this is normal, we don't care" response because they are big enough to know that eg podcasters need them.
I run pinecast.com. If there was a leaderboard for hn users serving XML, I'd almost certainly be in the top five. I don't disagree with your post. But: RSS downloads are at an all time low, and that's a bad thing. They're at an all time low because Spotify and Apple both fetch feeds from centralized servers. 1000 subscribers no longer means 24000ish daily feed fetches, it means 48. With keep alive or H2, these servic…
Also I agree that the re-centralisation is a bad thing, mainly.
(I'd like to move to email to discuss this further, if possible: I have an arXiv paper to write!)
Re: AI crawlers need to be more respectful
#113Earlier quoted context omitted.
I mean, blocking random IP ranges is your prerogative if you don't want to have customers. The scrapers will find ways around, while actual users will be unable to use your site. Residential proxies are something like $5 for 1000.
True, note domestic ISP IP ranges are published, and unless one deals internationally... don't bother serving people that will never buy anything from your firm anyways. Domestic "Users" functioning as proxies will be tripping usage limits, and getting temporarily banned. Google does this by the way, try hammering their services and find out what happens. Context cookies also immediately flag egregious multi-user rou…
And again, none of this is simple. It has taken me a few weeks to establish a usage mechanism that does catch the worst feed pullers, but it still can hurt legit new users. That is an opportunity cost.
Re: AI crawlers need to be more respectful
#114Earlier quoted context omitted.
Naming names isn't really required. The hosts have a $5,000 bandwidth fee, but so do the consumers. There's maybe 10 companies with the financial & compute resources to let a $5,000-per-month-per-website bug run rampant before taking the harvesting service offline. Meta/Google/Whoever may benefit from economies of scale, so they're not seeing the full $5,000 their side, but they're hitting tens of thousands of sites…
You know you can hit the data rate they were complaining about by using a residential fiber connection, right? 10 TB per day is about 1 Gigabit continuous if I’m not mistaken. There are probably millions of people that could to this if they wanted to.
Re: AI crawlers need to be more respectful
#115Invoice the abusers.
They're rolling in investor hype money, and they're obviously not spending it on competent developers if their bots behave like this, so there should be plenty left to cover costs.
Re: AI crawlers need to be more respectful
#116Earlier quoted context omitted.
True, note domestic ISP IP ranges are published, and unless one deals internationally... don't bother serving people that will never buy anything from your firm anyways. Domestic "Users" functioning as proxies will be tripping usage limits, and getting temporarily banned. Google does this by the way, try hammering their services and find out what happens. Context cookies also immediately flag egregious multi-user rou…
Those of us running sites for public information rather than sales cannot make the simple cut-off that you do. And again, none of this is simple. It has taken me a few weeks to establish a usage mechanism that does catch the worst feed pullers, but it still can hurt legit new users. That is an opportunity cost.
Allowing users known to have an active RAT or their "proxy friends" on a commercial site is not helping anyone.... especially the victims.
https://www.youtube.com/watch?v=aCbfMkh940Q
Worth studying the problem from time to time when you get bored of the antics.
These folks are generally uninterested in positively contributing to any community, but rather show up to cause trouble for fun and profit.
User API quotas are popular for a reason. =3
Re: AI crawlers need to be more respectful
#117Earlier quoted context omitted.
I mean, blocking random IP ranges is your prerogative if you don't want to have customers. The scrapers will find ways around, while actual users will be unable to use your site. Residential proxies are something like $5 for 1000.
Exactly, I'm banned or captchaed from half of all web sites these days, because of the AI.
Have a nice day, =3
Re: AI crawlers need to be more respectful
#118Earlier quoted context omitted.
I run pinecast.com. If there was a leaderboard for hn users serving XML, I'd almost certainly be in the top five. I don't disagree with your post. But: RSS downloads are at an all time low, and that's a bad thing. They're at an all time low because Spotify and Apple both fetch feeds from centralized servers. 1000 subscribers no longer means 24000ish daily feed fetches, it means 48. With keep alive or H2, these servic…
Do you have a CDN between you and Apple / Spotify? Because if you do I think that Apple/Spotify are polling that CDN every few minutes and the CDN is having its bandwidth wasted invisibly, but presumably priced in. Also I agree that the re-centralisation is a bad thing, mainly. (I'd like to move to email to discuss this further, if possible: I have an arXiv paper to write!)
Re: AI crawlers need to be more respectful
#119Earlier quoted context omitted.
Why are you ending all your messages with =3 ?
https://www.jpl.nasa.gov/images/pia22092-arp-142-the-penguin... Don't worry about it friend =3
Re: AI crawlers need to be more respectful
#120Earlier quoted context omitted.
https://www.jpl.nasa.gov/images/pia22092-arp-142-the-penguin... Don't worry about it friend =3
Read that before, read it again to make sure I didn't miss anything. The lack of clarity here is disappointing and only asks more questions than it answers.