Live data from Hacker News

LLM scraper bots are overloading acme.com's HTTPS server

acme.com

51–60 of 66 posts

Re: LLM scraper bots are overloading acme.com's HTTPS server

#51
post #3

> The LLM companies are not picking on me in particular, they are pounding every site on the net. Why is not this a criminal offense? They are hurting business for profit (or for higher valuation as they probably have no profit at all). Why are corporations allowed to do with impunity what could land even a teenager years in prison? Is there no rule of law anymore? The five-year and ten-year penalties kick in only wh…

His robots.txt explicitly allows bots including LLM bots to scrape his site

Re: LLM scraper bots are overloading acme.com's HTTPS server

#52
post #31

Earlier quoted context omitted.

A small part. On my server AI bots outnumber real visitors 300 to one.

How are you measuring this? Does your solution rely on user agent or device fingerprinting? Curious to know what tools are available today and how accurate they are.

I'm popular in Europe, there's no reason people from Singapore, Russia, Brazil and literally every other country in the world to all start visiting very old articles and permalinks for comments en masse.

Having honeypot links is the only thing that helps, but I'm running into massive IP tables, slowing things down.

This is not what I want to do with my time. I can't afford the expensive specialised tools. I'm just a solo entrepreneur on a shoestring budget. I just want to improve the website for my 3k real users and 10k real daily guests, not for bots.

Re: LLM scraper bots are overloading acme.com's HTTPS server

#53

Bot traffic is crazy even for smaller sites, but still manageable. I was getting 2,000 visitors a day on my infrequently updated website, but after I blocked all the bots via Cloudflare it went back to the normal double digit visitor count.

One day last week one of my clients' sites was getting about 2k "visitors" per second - I had to block the entire AS45102 to make it stop.

Re: LLM scraper bots are overloading acme.com's HTTPS server

#54
post #51
post #3

> The LLM companies are not picking on me in particular, they are pounding every site on the net. Why is not this a criminal offense? They are hurting business for profit (or for higher valuation as they probably have no profit at all). Why are corporations allowed to do with impunity what could land even a teenager years in prison? Is there no rule of law anymore? The five-year and ten-year penalties kick in only wh…

His robots.txt explicitly allows bots including LLM bots to scrape his site

The LLM scraper bots ignore robots.txt

Re: LLM scraper bots are overloading acme.com's HTTPS server

#55

Earlier quoted context omitted.

A small part. On my server AI bots outnumber real visitors 300 to one.

That such an absolutely ludicrous thing to hear in a "wtf are these people doing" type of way. I can't imagine a non-social media site would be generating enough traffic to the level that these bots need to be essentially doing continuous scraping. It's just gross to me to be okay with that level of unsophisticated effort that they just do the same thing over and over with zero gain.

Next to the massive amounts of energy they are burning in their own datacenters, they are burning up other datacenters as well. Plus all the extra energy used by every router, hub and switch in between.

Re: LLM scraper bots are overloading acme.com's HTTPS server

#56
post #50

> I closed port 443 > Now closing https service is obviously just a temporary fix Probably the best starting point would be to edit the robots.txt file and disallow LLM bots there. Currently the file allows all bots: http://acme.com/robots.txt

The LLM scraper bots ignore robots.txt

Re: LLM scraper bots are overloading acme.com's HTTPS server

#57
post #40

Earlier quoted context omitted.

I don't mean that users are following the links to `acme.com` and `demo.com` type domains in documentation; I mean that bots are likely finding and following many links to them because of their widespread use in documentation. If you search for `site:github.com "acme.com"` in Google, you'll find numerous instances of the domain being used in contrived links in documentation as an example of how URLs might be structur…

That is very possible . But it is not necessary to see the results that are being described. If sites like my tiny little browser game, with roughly 120 weekly unique users, are getting absolutely hammered by the scraper-bots (it was, last year, until I put the Wiki behind a login wall; now I still get a significant amount of bot traffic, it's just no longer enough to actually crash the game), then sites that people…

The article describes that a lot of the requests are for non-existent URLs. Do you observe the same?

Re: LLM scraper bots are overloading acme.com's HTTPS server

#58
post #50

> I closed port 443 > Now closing https service is obviously just a temporary fix Probably the best starting point would be to edit the robots.txt file and disallow LLM bots there. Currently the file allows all bots: http://acme.com/robots.txt

[dead]

Re: LLM scraper bots are overloading acme.com's HTTPS server

#59
post #3

> The LLM companies are not picking on me in particular, they are pounding every site on the net. Why is not this a criminal offense? They are hurting business for profit (or for higher valuation as they probably have no profit at all). Why are corporations allowed to do with impunity what could land even a teenager years in prison? Is there no rule of law anymore? The five-year and ten-year penalties kick in only wh…

adapt or die waiting on the govt to do something is a path of failure

> waiting on the govt to do something is a path of failure

To keep the goverment accountable is a duty of every citizen and the only way to have a functioning society. The failure is to let the goverment be arbitrary and cater to the powerful instead of following the rule of law and applying it equally at all levels.

Re: LLM scraper bots are overloading acme.com's HTTPS server

#60
post #12

Earlier quoted context omitted.

Is what an offence lol? Bot scraper traffic? How do you think search engines work?

They work because they offer ways to opt out, they honor crawl delay, setting ideal scraping times, IndexNow, etc. And they give you real, valuable traffic in return.

Most offer ways to opt out, some don’t. Scraping somebody’s website might be annoying or problematic traffic-wise but that’s a far (very far) step removed from saying scrapers should be criminalised. The latter statement is outright laughable.
Post reply on HN