Live data from Hacker News

Amazon's AI crawler is making my Git server unstable

xeiaso.net

81–90 of 261 posts

Re: Amazon's AI crawler is making my Git server unstable

#81
post #74
post #72

Earlier quoted context omitted.

In the UK, the Computer Misuse Act applies if: * There is knowledge that the intended access was unauthorised * There is an intention to secure access to any program or data held in a computer I imagine US law has similar definitions of unauthorized access? `robots.txt` is the universal standard for defining what is unauthorised access for bots. No programmer could argue they aren't aware of this, and ignoring it, fo…

> `robots.txt` is the universal standard Quite the assumption, you just upset a bunch of alien species.

Dammit. Unchecked geocentric model privilege, sorry about that.

Re: Amazon's AI crawler is making my Git server unstable

#82

Earlier quoted context omitted.

You should check your websites like grass dot io (I refuse to give them traffic). They pay you for your bandwidth while they resell it to 3rd parties, which is why a lot of bot traffic looks like it comes from residential IPs.

Yes, but the point is that big company crawlers aren’t paying for questionably sourced residential proxies. If this person is seeing a lot of traffic from residential IPs then I would be shocked if it’s really Amazon. I think someone else is doing something sketchy and they put “AmazonBot” in the user agent to make victims think it’s Amazon. You can set the user agent string to anything you want, as we all know.

> Yes, but the point is that big company crawlers aren’t paying for questionably sourced residential proxies

You'd be surprised...

Re: Amazon's AI crawler is making my Git server unstable

#83

Earlier quoted context omitted.

You should check your websites like grass dot io (I refuse to give them traffic). They pay you for your bandwidth while they resell it to 3rd parties, which is why a lot of bot traffic looks like it comes from residential IPs.

Yes, but the point is that big company crawlers aren’t paying for questionably sourced residential proxies. If this person is seeing a lot of traffic from residential IPs then I would be shocked if it’s really Amazon. I think someone else is doing something sketchy and they put “AmazonBot” in the user agent to make victims think it’s Amazon. You can set the user agent string to anything you want, as we all know.

They could be using echo devices to proxy their traffic…

Although I’m not necessarily gonna make that accusation, because it would be pretty serious misconduct if it were true.

Re: Amazon's AI crawler is making my Git server unstable

#84

Earlier quoted context omitted.

You should check your websites like grass dot io (I refuse to give them traffic). They pay you for your bandwidth while they resell it to 3rd parties, which is why a lot of bot traffic looks like it comes from residential IPs.

Yes, but the point is that big company crawlers aren’t paying for questionably sourced residential proxies. If this person is seeing a lot of traffic from residential IPs then I would be shocked if it’s really Amazon. I think someone else is doing something sketchy and they put “AmazonBot” in the user agent to make victims think it’s Amazon. You can set the user agent string to anything you want, as we all know.

I worked for Microsoft doing malware detection back 10+ years ago, and questionably sourced proxies were well and truly on the table

Re: Amazon's AI crawler is making my Git server unstable

#86
post #3

Upvoted because we’re seeing the same behavior from all AI and Seo bots. They’re BARELY respecting Robots.txt, and hard to block. And when they crawl, they spam and drive up load so high they crash many servers for our clients. If AI crawlers want access they can either behave, or pay. The consequence will almost universal blocks otherwise!

> The consequence will almost universal blocks otherwise! Who cares? They've already scraped the content by then.

If they only needed a one-time scrape we really wouldn't be seeing noticeable not traffic today.

Re: Amazon's AI crawler is making my Git server unstable

#89
post #45

Earlier quoted context omitted.

This submission has nothing to do with IP laundering. The bot is straining their server and causing OP technical issues.

Commentary is often second- and third-order.

True, but it tends to flow there organically. This comment was off topic from the start.

Re: Amazon's AI crawler is making my Git server unstable

#90

I don’t think I’d assume this is actually Amazon. The author is seeing requests from rotating residential IPs and changing user agent strings > It's futile to block AI crawler bots because they lie, change their user agent, use residential IP addresses as proxies, and more. Impersonating crawlers from big companies is a common technique for people trying to blend in. The fact that requests are coming from residential…

I wouldn't put it past any company these days doing crawling in an aggressive manner to use proxy networks.
Post reply on HN