I don't look at the logs of my personal site very often as it is a static site, but just went to check, and yeah, it's almost all ai crawlers. Note sure what is going on here, but I hope this isn't really anthropic: 34.124.XXX.XXX - - [07/Aug/2026:07:55:31 -0700] "GET /secrets.json HTTP/2.0" 404 343 "-" "anthropic-ai" 34.124.XXX.XXX - - [07/Aug/2026:07:55:31 -0700] "GET /credentials.json HTTP/2.0" 404 343 "-" "anthro…
A year of fighting scrapers on my 1.5 million-page website
301–310 of 451 posts
Re: A year of fighting scrapers on my 1.5 million-page website
#302Earlier quoted context omitted.
I wish they'd ease up on the whole "don't change the logo without paying us" thing. The furry anime character is a turnoff for anyone with a brand or personal image that doesn't mesh with those subcultures.
It's actually a genius idea. If you are someone who the professional presentation of not having an anime girl on the loading page is required, then you can afford to fork over the cash to fund development.
Which one then wonders why they would install it in the first place.
Re: A year of fighting scrapers on my 1.5 million-page website
#303The worrying thing here is that so many people accepted outsourcing the decision on who can see their website to a large company (Cloudflare). If the company decides that a certain user should not see the website, the user will not see the website, and no one will know about it, and the user will have no recourse. That is not the open web that I would like to see. A second side effect of a knee-jerk reaction to bots…
Re: A year of fighting scrapers on my 1.5 million-page website
#304Earlier quoted context omitted.
This is such a grim thing to read.
I disagree. I personally consider these my biggest problems with the Web: - Bias, specifically commercial bias - Webpage formatting: every website looks different, hides the information I want in different places. Sure, it "looks pretty", but I don't want pretty; I want info - Scams/SEO/etc... The LLMs are very good at reading from multiple sources, parsing, and presenting only the data in a consistent format. There…
So why should I pick an LLM over google's inbuilt LLM or an actual cookbook on the desk? What wasted time is the LLM saving?
Re: A year of fighting scrapers on my 1.5 million-page website
#305The worrying thing here is that so many people accepted outsourcing the decision on who can see their website to a large company (Cloudflare). If the company decides that a certain user should not see the website, the user will not see the website, and no one will know about it, and the user will have no recourse. That is not the open web that I would like to see. A second side effect of a knee-jerk reaction to bots…
I have an old os (macos 10.11), running the highest firefox esr I can run, and I get denied by cloudflare.
But not always immediately - I get to enable javascript/cookies sometimes just to be denied.
they are not the folks we want gatekeeping the internet, they are opportunists increasing their OKRs
Re: A year of fighting scrapers on my 1.5 million-page website
#306Re: A year of fighting scrapers on my 1.5 million-page website
#307Earlier quoted context omitted.
What if I’m a web alerts company? I crawl your content but my clients are all end users who actually see your content at your site.
Then you are a uninvited bot that crawls my content in order to send alarms to people to tell them to visit my site? This is a "what if it's for your own good?" argument. What if I break into your house to clean your toilets? Although amusingly, if you're a "web alerts company" that I didn't contract and sends out alerts in batches on your own schedule, you will probably send a ton of your customers to my site at the…
Re: A year of fighting scrapers on my 1.5 million-page website
#308Earlier quoted context omitted.
> You take your site off the public internet and paywall it off I did that. I even printed and bound it in a bunch of dead trees and put it up for sale. The LLMs still stole it.
Paywalling off the site solves the problem of monetizing bot traffic, any bot crawling your pages paid you to be there, but it can't fix the plagiarism/copyright infringement problem
Re: A year of fighting scrapers on my 1.5 million-page website
#309at this point the web is mostly machines politely asking other machines for permission to read each other's content, and humans are the ones triggering the captchas.
Re: A year of fighting scrapers on my 1.5 million-page website
#310Anubis[1] is a superb fix for sites not behind Cloudflare/Fastly/Bunny etc. We had millions of bot requests, on a site serving all countries so we couldn't block by country, with fake user-agents so we couldn't block using that. It uses 'proof of work' to detect real browser software. [1] https://anubis.techaro.lol/
I wish they'd ease up on the whole "don't change the logo without paying us" thing. The furry anime character is a turnoff for anyone with a brand or personal image that doesn't mesh with those subcultures.
If you can't be bothered to maintain a trivial fork, then why should the author of anubis be bothered to serve your branding needs? It's not like you have a service contract or anything do you?