Live data from Hacker News

Amazon's AI crawler is making my Git server unstable

xeiaso.net

31–40 of 261 posts

Re: Amazon's AI crawler is making my Git server unstable

#31
post #26

[flagged]

I think there’s a difference between crawling websites at a reasonable pace instead of just hammering the server to the point it’s unusable.

Nobody has problems with the Google Search indexer trying to crawl websites in a responsible way

Re: Amazon's AI crawler is making my Git server unstable

#33
post #3

Upvoted because we’re seeing the same behavior from all AI and Seo bots. They’re BARELY respecting Robots.txt, and hard to block. And when they crawl, they spam and drive up load so high they crash many servers for our clients. If AI crawlers want access they can either behave, or pay. The consequence will almost universal blocks otherwise!

> The consequence will almost universal blocks otherwise! Who cares? They've already scraped the content by then.

Bold to assume that an AI scraper won't come back to download everything again, just in case there's any new scraps of data to extract. OP mentioned in the other thread that this bot had pulled 3TB so far, and I doubt their git server actually has 3TB of unique data, so the bot is probably pulling the same data over and over again.

Re: Amazon's AI crawler is making my Git server unstable

#34
post #26

[flagged]

Most of those of those artists aren’t any better though. I’m on a couple artists’ forums and outlets like Tumblr, and I saw firsthand the immediate, total 180 re: IP protection when genAI showed up. Overnight, everybody went from “copying isn’t theft, it leaves the original!” and other such mantras, to being die-hard IP maximalists. To say nothing of how they went from “anything can be art and it doesn’t matter what tools you’re using” to forming witch-hunt mobs against people suspected of using AI tooling. AI has made a hypocrite out of everybody.

Re: Amazon's AI crawler is making my Git server unstable

#35
Are you sure it isn't a DDoS masquerading as Amazon?

Requests coming from residential ips is really suspicious.

Edit: the motivation for such a DDoS might be targeting Amazon, by taking down smaller sites and making it look like amazon is responsible.

If it is Amazon one place to start is blocking all the the ip ranges they publish. Although it sounds like there are requests outside those ranges...

Re: Amazon's AI crawler is making my Git server unstable

#36
post #26

[flagged]

I think there’s a difference between crawling websites at a reasonable pace instead of just hammering the server to the point it’s unusable. Nobody has problems with the Google Search indexer trying to crawl websites in a responsible way

For sure.

I'm really just pointing out the inconsistent technocrat attitude towards labor, sovereignty, and resources.

Re: Amazon's AI crawler is making my Git server unstable

#37
post #23

Can demonstrable ignoring of robots.txt help the cases of copyright infringement lawsuits against the "AI" companies, their partners, and customers?

On what legal basis?

Terms of use contract violation?

Re: Amazon's AI crawler is making my Git server unstable

#39
post #13
post #8

I had this same issue recently. My Forgejo instance started to use 100 % of my home server's CPU as Claude and its AI friends from Meta and Google were hitting the basically infinite links at a high rate. I managed to curtail it with robots.txt and a user agent based blocklist in Caddy, but who knows how long that will work. Whatever happened to courtesy in scraping?

> Whatever happened to courtesy in scraping? Money happened. AI companies are financially incentivized to take as much data as possible, as quickly as possible, from anywhere they can get it, and for now they have so much cash to burn that they don't really need to be efficient about it.

Need to act fast before the copyright cases in the court gets handled.
Post reply on HN