Live data from Hacker News

Tell HN: Amazonbot aggressively scraping my website and ignoring robots.txt

news.ycombinator.com

1–10 of 26 posts

Tell HN: Amazonbot aggressively scraping my website and ignoring robots.txt

#1
At the beginning of the year I decided to set up a scraping and LLM honeypot on one of my personal websites which included a fake git repo with code containing fake HTTP endpoints. The address to this repo was hidden in a public page inside a comment.

About three weeks ago IP addresses from Amazon Searchbot attempted to make requests to the fake endpoints included inside a shell script.

My robots.txt explicitly includes Amazonbot.

I am honestly surprised that this is coming from Amazon. Is this kind of behavior legal?

Re: Tell HN: Amazonbot aggressively scraping my website and ignoring robots.txt

#3
Respecting robots.txt is not a legal obligation. If you don’t want the traffic, ban the offending IP block or ASN, or put the origin behind Cloudflare for their crawler blocker.

You could go down the rabbit hole and try to find someone at Amazon to tell their crawler to chill. Don’t expect success but would make a fun blog post.

https://www.cloudflare.com/learning/ai/how-to-block-ai-crawl...

Re: Tell HN: Amazonbot aggressively scraping my website and ignoring robots.txt

#4

Respecting robots.txt is not a legal obligation. If you don’t want the traffic, ban the offending IP block or ASN, or put the origin behind Cloudflare for their crawler blocker. You could go down the rabbit hole and try to find someone at Amazon to tell their crawler to chill. Don’t expect success but would make a fun blog post. https://www.cloudflare.com/learning/ai/how-to-block-ai-crawl...

My concern is more in relation to trying to use an endpoint from non-public source code, it seems negligent to randomly try endpoints like this

Re: Tell HN: Amazonbot aggressively scraping my website and ignoring robots.txt

#5
post #4

Respecting robots.txt is not a legal obligation. If you don’t want the traffic, ban the offending IP block or ASN, or put the origin behind Cloudflare for their crawler blocker. You could go down the rabbit hole and try to find someone at Amazon to tell their crawler to chill. Don’t expect success but would make a fun blog post. https://www.cloudflare.com/learning/ai/how-to-block-ai-crawl...

My concern is more in relation to trying to use an endpoint from non-public source code, it seems negligent to randomly try endpoints like this

Whether it’s negligent isn’t terribly relevant. Is it illegal? It is not, they aren’t bypassing access controls. No different than using Shodan, scanning public IPs, crawling open directories, etc.

It would be different if they were attempting to brute force credentials to access an endpoint, but they aren’t.

Re: Tell HN: Amazonbot aggressively scraping my website and ignoring robots.txt

#6
Assuming your site has no dependency on Amazon EC2:

    fetch_aws()
    {
     curl -s \
     -H "Accept-Charset: UTF-8" \
     -H "Accept-Encoding: gzip, deflate" \
     -H "Accept: text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8" \
     -H "Accept-Language: en" \
     -H "User-Agent: Mozilla/5.0 (Windows NT 11.1; rv:102.0; ) Gecko/20100101" \
     -H "Cache-Control: 0" \
     -H "Host: ip-ranges.amazonaws.com" \
     -H "Connection: keep-alive" \
     --referer https://ec2.amazon.com/ \
     -o /dev/shm/aws.json \
     --url "https://ip-ranges.amazonaws.com/ip-ranges.json"
    grep ip_prefix /dev/shm/aws.json | awk -F "\"" '{print $4}' | sort -n | uniq
    rm -f /dev/shm/aws.json;
    };
Make sure that gives you CIDR blocks, then add it to a script and use a for loop to

    for AWS in $(fetch_aws);do ip route add blackhole "${AWS}" 2>/dev/null

Google and Cloudflare if you need them:

    get_google()
    {
    for line in $(dig +short txt _cloud-netblocks.googleusercontent.com | tr " " "\n" | grep include | cut -f 2 -d :)                                                           
    do                                                                                                                                                                          
            dig +short txt "${line}"                                                                                                                                            
    done | tr " " "\n" | grep ip4 | cut -f 2 -d : | sort -n | uniq
    }
    
    get_cloudflare()
    {
    curl -A Mozilla "https://api.cloudflare.com/local-ip-ranges.csv"|grep -Ev "::|/32"|awk -F "," '{print $1}'|sort | uniq
    }
This assumes you have no need to connect to or get connections from Amazon on your web server. If your server is an instance in Amazon that will break DNS resolution unless you are using something outside of AWS for DNS. The gateway should still work just fine. Obviously test from an out of band console if that is an option. This is easier to maintain if you remove DNS records for IPv6 and eventually disable IPv6 listeners.

If you paste a couple of lines of the bots from your access logs I can offer more suggestions in the event they try from outside of Amazon.

If you want to have some fun, add a hidden link only the bot will see that points to http://cpanel.yourdomain.tld/ after adding a DNS record for cpanel that points to 169.254.169.254 so they start scraping the AWS cloud-init IP.

Re: Tell HN: Amazonbot aggressively scraping my website and ignoring robots.txt

#8
post #6

Assuming your site has no dependency on Amazon EC2: fetch_aws() { curl -s \ -H "Accept-Charset: UTF-8" \ -H "Accept-Encoding: gzip, deflate" \ -H "Accept: text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8" \ -H "Accept-Language: en" \ -H "User-Agent: Mozilla/5.0 (Windows NT 11.1; rv:102.0; ) Gecko/20100101" \ -H "Cache-Control: 0" \ -H "Host: ip-ranges.amazonaws.com" \ -H "Connection: keep-alive" \ --re…

Thanks, the Amazonbot IPs that accessed my honeypot are in this other list:

https://developer.amazon.com/amazonbot/searchbot-ip-addresse...

By the way, they used false user-agents.

Re: Tell HN: Amazonbot aggressively scraping my website and ignoring robots.txt

#9
I haven't personally seen it, but I think us operators need to take entire registered ranges and hook 'em up to fail2ban. Maybe maintain lists of offenders and document it and block these orgs.

That is, I haven't personally seen a bad actors list that gets used in a fail2ban-like setup.

Post reply on HN