Live data from Hacker News

A year of fighting scrapers on my 1.5 million-page website

patronview.com

301–310 of 451 posts

Re: A year of fighting scrapers on my 1.5 million-page website

#301
post #230

I don't look at the logs of my personal site very often as it is a static site, but just went to check, and yeah, it's almost all ai crawlers. Note sure what is going on here, but I hope this isn't really anthropic: 34.124.XXX.XXX - - [07/Aug/2026:07:55:31 -0700] "GET /secrets.json HTTP/2.0" 404 343 "-" "anthropic-ai" 34.124.XXX.XXX - - [07/Aug/2026:07:55:31 -0700] "GET /credentials.json HTTP/2.0" 404 343 "-" "anthro…

so they broke into your server?

Re: A year of fighting scrapers on my 1.5 million-page website

#302

Earlier quoted context omitted.

I wish they'd ease up on the whole "don't change the logo without paying us" thing. The furry anime character is a turnoff for anyone with a brand or personal image that doesn't mesh with those subcultures.

It's actually a genius idea. If you are someone who the professional presentation of not having an anime girl on the loading page is required, then you can afford to fork over the cash to fund development.

I guess the disconnect here is a bunch of HN'ers believing professional companies and websites want to attach their branding to a sexualized anime character and that they are willing to pay to remove it.

Which one then wonders why they would install it in the first place.

Re: A year of fighting scrapers on my 1.5 million-page website

#303
post #101

The worrying thing here is that so many people accepted outsourcing the decision on who can see their website to a large company (Cloudflare). If the company decides that a certain user should not see the website, the user will not see the website, and no one will know about it, and the user will have no recourse. That is not the open web that I would like to see. A second side effect of a knee-jerk reaction to bots…

Jokes on them, the second I see that “Verifying you are human…” redirect I leave and never return.

Re: A year of fighting scrapers on my 1.5 million-page website

#304

Earlier quoted context omitted.

This is such a grim thing to read.

I disagree. I personally consider these my biggest problems with the Web: - Bias, specifically commercial bias - Webpage formatting: every website looks different, hides the information I want in different places. Sure, it "looks pretty", but I don't want pretty; I want info - Scams/SEO/etc... The LLMs are very good at reading from multiple sources, parsing, and presenting only the data in a consistent format. There…

Why even use a LLM, just buy a cookbook and you can even waste less time and know the information is valid. And again if your logic is consistent, google routes your search to an LLM automatically and renders the reply - that will save as much time and do the same thing since LLMs are so good at parsing and presenting information.

So why should I pick an LLM over google's inbuilt LLM or an actual cookbook on the desk? What wasted time is the LLM saving?

Re: A year of fighting scrapers on my 1.5 million-page website

#305
post #101

The worrying thing here is that so many people accepted outsourcing the decision on who can see their website to a large company (Cloudflare). If the company decides that a certain user should not see the website, the user will not see the website, and no one will know about it, and the user will have no recourse. That is not the open web that I would like to see. A second side effect of a knee-jerk reaction to bots…

I am denied by cloudflare CONSTANTLY on one system.

I have an old os (macos 10.11), running the highest firefox esr I can run, and I get denied by cloudflare.

But not always immediately - I get to enable javascript/cookies sometimes just to be denied.

they are not the folks we want gatekeeping the internet, they are opportunists increasing their OKRs

Re: A year of fighting scrapers on my 1.5 million-page website

#306
post #188

Earlier quoted context omitted.

What if I’m a web alerts company? I crawl your content but my clients are all end users who actually see your content at your site.

Fetch my RSS then.

The vast majority of news sites don't have an RSS. They don't even have a robots.txt.

Re: A year of fighting scrapers on my 1.5 million-page website

#307

Earlier quoted context omitted.

What if I’m a web alerts company? I crawl your content but my clients are all end users who actually see your content at your site.

Then you are a uninvited bot that crawls my content in order to send alarms to people to tell them to visit my site? This is a "what if it's for your own good?" argument. What if I break into your house to clean your toilets? Although amusingly, if you're a "web alerts company" that I didn't contract and sends out alerts in batches on your own schedule, you will probably send a ton of your customers to my site at the…

No, I won't send a ton, I'll send a few dozens to a few hundreds at most because not everyone is interested in the same things. And they'll visit at their own timeline. If your site can't service a few dozen requests simultaneously then you probably aren't a news site in the first place so the whole argument is moot.

Re: A year of fighting scrapers on my 1.5 million-page website

#308

Earlier quoted context omitted.

> You take your site off the public internet and paywall it off I did that. I even printed and bound it in a bunch of dead trees and put it up for sale. The LLMs still stole it.

Paywalling off the site solves the problem of monetizing bot traffic, any bot crawling your pages paid you to be there, but it can't fix the plagiarism/copyright infringement problem

Well... They paid someone to be there, not necessarily you.

Re: A year of fighting scrapers on my 1.5 million-page website

#309

at this point the web is mostly machines politely asking other machines for permission to read each other's content, and humans are the ones triggering the captchas.

Machines don't pay my bills though. If the content was free to share widely, then yea, who cares how it's accessed.

Re: A year of fighting scrapers on my 1.5 million-page website

#310

Anubis[1] is a superb fix for sites not behind Cloudflare/Fastly/Bunny etc. We had millions of bot requests, on a site serving all countries so we couldn't block by country, with fake user-agents so we couldn't block using that. It uses 'proof of work' to detect real browser software. [1] https://anubis.techaro.lol/

I wish they'd ease up on the whole "don't change the logo without paying us" thing. The furry anime character is a turnoff for anyone with a brand or personal image that doesn't mesh with those subcultures.

How hard is it to maintain a fork that changes nothing except the logo?

If you can't be bothered to maintain a trivial fork, then why should the author of anubis be bothered to serve your branding needs? It's not like you have a service contract or anything do you?

Post reply on HN