Live data from Hacker News

Show HN: Stop AI scrapers from hammering your self-hosted blog (using porn)

github.com

21–30 of 288 posts

Re: Show HN: Stop AI scrapers from hammering your self-hosted blog (using porn)

#21

I own a forum which currently has 23k online users, all of them bots. The last new post in that forum is from _2019_. Its topic is also very niche. Why are so many bots there? This site should have basically been scraped a million times by now, yet those bots seem to fetch the stuff live, on the fly? I don’t get it.

How do you define a user, and how do you define online?

If the forum considers unique cookies to be a user and creates a new cookie for any new cookie-less request, and if it considers a user to be online for 1 hour after their last request, then actually this may be one scraper making ~6 requests per second. That may be a pain in its own way, but it's far from 23k online bots.

Re: Show HN: Stop AI scrapers from hammering your self-hosted blog (using porn)

#22
post #4

Nice! Reminds me of “Piracy as Proof of Personhood”. If you want to read that one go to Paged Out magazine (at https://pagedout.institute/ ), navigate to issue #7, and flip to page 9. I wonder if this will start making porn websites rank higher in google if it catches on… Have you tested it with the Lynx web browser? I bet all the links would show up if a user used it. Oh also couldn’t AI scrapers just start imperson…

> Hey I wonder if there is some situation where negative SEO would be a good tactic. Generally though I think if you wanted something to stay hidden it just shouldn’t be on a public web server.

At least once upon a time there was a pirate textbook library that used HTTP basic auth with a prompt that made the password really easy to guess. I suppose the main goal was to keep crawlers out even if they don't obey robots.txt, and at the same time be as easy for humans as possible.

Re: Show HN: Stop AI scrapers from hammering your self-hosted blog (using porn)

#24

I own a forum which currently has 23k online users, all of them bots. The last new post in that forum is from _2019_. Its topic is also very niche. Why are so many bots there? This site should have basically been scraped a million times by now, yet those bots seem to fetch the stuff live, on the fly? I don’t get it.

When you have trillions of dollars being poured into your company by the financial system, and when furthermore there are no repercussions for behaving however you please, you tend not to care about that sort of "waste".

Re: Show HN: Stop AI scrapers from hammering your self-hosted blog (using porn)

#28
post #22
post #4

Nice! Reminds me of “Piracy as Proof of Personhood”. If you want to read that one go to Paged Out magazine (at https://pagedout.institute/ ), navigate to issue #7, and flip to page 9. I wonder if this will start making porn websites rank higher in google if it catches on… Have you tested it with the Lynx web browser? I bet all the links would show up if a user used it. Oh also couldn’t AI scrapers just start imperson…

> Hey I wonder if there is some situation where negative SEO would be a good tactic. Generally though I think if you wanted something to stay hidden it just shouldn’t be on a public web server. At least once upon a time there was a pirate textbook library that used HTTP basic auth with a prompt that made the password really easy to guess. I suppose the main goal was to keep crawlers out even if they don't obey robots…

Interesting note, thank you.

Re: Show HN: Stop AI scrapers from hammering your self-hosted blog (using porn)

#29

Cloudflare offers bot mitigation for free, and pretty generous WAF rules that makes mitigations like this seem a little overblown to me

For “free”.

Did you put “free” in quotes because you need to have paid for stuff from cloudflare to use the “free” thing?

If so, I suppose it’s like those magazines that say ”free cd”.

Re: Show HN: Stop AI scrapers from hammering your self-hosted blog (using porn)

#30
post #6

Earlier quoted context omitted.

hey! thanks for that read suggestion that's indeed a pretty funny captcha strat. Yup the links show up if you use the Lynx web browser. As for AI scrapers impersonating googlebot I feel like yes they'd definitely start doing that, unless the risk of getting sued by google is too high? If google could even sue them for doing that? Not an internet litigation expert but seems like it could be debatable

Yeah I guess I don’t know if you can sue someone for using your headers, would be interesting to see how that goes.

i think making the case of "you are acting (sending web requests) while knowingly identifying as another legal entity (and criminally/libelously/etc)" shouldn't be toooo hard
Post reply on HN