Live data from Hacker News

Amazon's AI crawler is making my Git server unstable

xeiaso.net

201–210 of 261 posts

Re: Amazon's AI crawler is making my Git server unstable

#201
post #3

Upvoted because we’re seeing the same behavior from all AI and Seo bots. They’re BARELY respecting Robots.txt, and hard to block. And when they crawl, they spam and drive up load so high they crash many servers for our clients. If AI crawlers want access they can either behave, or pay. The consequence will almost universal blocks otherwise!

Don't worry, though, because IP law only applies to peons like you and me. :)

Re: Amazon's AI crawler is making my Git server unstable

#202
post #52
post #3

Upvoted because we’re seeing the same behavior from all AI and Seo bots. They’re BARELY respecting Robots.txt, and hard to block. And when they crawl, they spam and drive up load so high they crash many servers for our clients. If AI crawlers want access they can either behave, or pay. The consequence will almost universal blocks otherwise!

Is there some way website can sell those Data to AI bot in a large zip file rather than being constantly DDoS? Or they could at least have the curtesy to scrap during night time / off peak hours.

Is existing intellectual property law not sufficient? Why aren't companies being prosecuted for large-scale theft?

Re: Amazon's AI crawler is making my Git server unstable

#203
post #3

Upvoted because we’re seeing the same behavior from all AI and Seo bots. They’re BARELY respecting Robots.txt, and hard to block. And when they crawl, they spam and drive up load so high they crash many servers for our clients. If AI crawlers want access they can either behave, or pay. The consequence will almost universal blocks otherwise!

> The consequence will almost universal blocks otherwise! Who cares? They've already scraped the content by then.

That's the spirit!

Re: Amazon's AI crawler is making my Git server unstable

#204
post #3

Upvoted because we’re seeing the same behavior from all AI and Seo bots. They’re BARELY respecting Robots.txt, and hard to block. And when they crawl, they spam and drive up load so high they crash many servers for our clients. If AI crawlers want access they can either behave, or pay. The consequence will almost universal blocks otherwise!

Global tarpit is the solution. It makes sense anyway even without taking AI crawlers into account. Back when I had to implement that, I went the semi manual route - parse the access log and any IP address averaging more than X hits a second on /api gets a -j TARPIT with iptables [1]. Not sure how to implement it in the cloud though, never had the need for that there yet. [1] https://gist.github.com/flaviovs/103a0dbf6…

Don't we have intellectual property law for this tho?

Re: Amazon's AI crawler is making my Git server unstable

#205
post #8

I had this same issue recently. My Forgejo instance started to use 100 % of my home server's CPU as Claude and its AI friends from Meta and Google were hitting the basically infinite links at a high rate. I managed to curtail it with robots.txt and a user agent based blocklist in Caddy, but who knows how long that will work. Whatever happened to courtesy in scraping?

When will the last Hacker News realize that Meta and OpenAI and every last massive tech company were always going to screw us all over for a quick buck?

Remember, Facebook famously made it easy to scrape your friends from MySpace, and then banned that exact same activity from their site once they got big.

Wake the f*ck up.

Re: Amazon's AI crawler is making my Git server unstable

#206

Earlier quoted context omitted.

Yes, but the point is that big company crawlers aren’t paying for questionably sourced residential proxies. If this person is seeing a lot of traffic from residential IPs then I would be shocked if it’s really Amazon. I think someone else is doing something sketchy and they put “AmazonBot” in the user agent to make victims think it’s Amazon. You can set the user agent string to anything you want, as we all know.

I wonder if anyone has checked whether Alexa devices serve as a private proxy network for AmazonBot’s use.

Yes, people have probably analyzed Alexa traffic once or twice over the years.

Re: Amazon's AI crawler is making my Git server unstable

#207

The best way to fight this would not to block them, that does not cause Amazon/others anything. (clearly). What if instead it was possible to feed the bots clearly damaging and harmfull content? If done on a larger scale, and Amazon discovers the poisoned pills they could have to spend money rooting it out, quick like, and make attempts to stop their bots to ingest it. Of course nobody wants to have that tuff on thei…

Why is it my responsibility to piss into the wind because these billionaire companies keep getting to break the law with impunity?

Re: Amazon's AI crawler is making my Git server unstable

#208
post #23

Can demonstrable ignoring of robots.txt help the cases of copyright infringement lawsuits against the "AI" companies, their partners, and customers?

On what legal basis?

I wind up in jail for ten years if I download an episode of iCarly; Sam Altman inhales every last byte on the internet and gets a ticker tape parade. Make it make sense.

Re: Amazon's AI crawler is making my Git server unstable

#209
post #145

Unless we start chopping these tech companies down there's not much hope for the public internet. They now have an incentive to crawl anything they can and have vastly more resources than even most governments. Most resources I need to host in a way that's internet facing are behind keyauth and I'm not sure I see a way around doing that for at least a while

That is an unworkable solution considering most of the people here are employed by these companies.

Re: Amazon's AI crawler is making my Git server unstable

#210

What are the actual rules/laws about scraping? I have a few projects I'd like to do that involve scraping but have always been conscious about respecting the host's servers, plus whether private content is copyrighted. But sounds like AI companies don't give a shit lol. If anyone has a good resource on the subject I'd be grateful!

If you go to a police station and ask them to arrest Amazon for accessing your website too often, will they arrest Amazon, or laugh at you? While facetious in nature, my point is that people walking around in real brick and mortar locations simply do not care. If you want police to enforce laws, those are the kinds of people that need to care about your problem. Until that occurs, youll have to work around the proble…

Oh, police will happily enforce IP law, if the MPAA or RIAA want them to. You only get a pass if you're a thought leader.
Post reply on HN