Live data from Hacker News

A year of fighting scrapers on my 1.5 million-page website

patronview.com

111–120 of 451 posts

Re: A year of fighting scrapers on my 1.5 million-page website

#111
post #101

The worrying thing here is that so many people accepted outsourcing the decision on who can see their website to a large company (Cloudflare). If the company decides that a certain user should not see the website, the user will not see the website, and no one will know about it, and the user will have no recourse. That is not the open web that I would like to see. A second side effect of a knee-jerk reaction to bots…

> I wanted the information and I have no way of getting at it without investing my own time and effort in going to your website.

I think you've hit the nail on the head here. I think that's a big reason why people want to ban bots.

Re: A year of fighting scrapers on my 1.5 million-page website

#112
post #10

I'm in a similar boat. Probably 99.999% is bots. I have nearly 1m unique "visitors" according to cloudflare and my real users are in the dozens a day. That being said, I love the open internet and am holding on to keeping as much open as I can.

But what do these bots gain from this?

Re: A year of fighting scrapers on my 1.5 million-page website

#113
post #101

The worrying thing here is that so many people accepted outsourcing the decision on who can see their website to a large company (Cloudflare). If the company decides that a certain user should not see the website, the user will not see the website, and no one will know about it, and the user will have no recourse. That is not the open web that I would like to see. A second side effect of a knee-jerk reaction to bots…

>That is not the open web that I would like to see.

Cloudflare is opt-in so I don't see that being an issue (yet).

Re: A year of fighting scrapers on my 1.5 million-page website

#114
post #103
post #68

Earlier quoted context omitted.

There's a difference between someone running a scraping tool occasionally and bots constantly and rapidly re-scraping the same site over and over again

This is an important context: the site is more likely to be targeted by scrapers because it is a curated collection of scraped information.

Do the bots care? Seemingly very little intelligence in many of them. Could be as simple as the site has more pages, so more traffic.

Loot first, ask questions later.

Re: A year of fighting scrapers on my 1.5 million-page website

#115
post #5

Earlier quoted context omitted.

Because the ubermensch on HN can detect AI text from 10 miles away with 100% accuracy... like that time they called documents from 2015 AI written.

I'm just beyond thrilled that it's actually starting to be a public embarrassment to cough up a glob of AI text.

[deleted]

Re: A year of fighting scrapers on my 1.5 million-page website

#116
post #101

The worrying thing here is that so many people accepted outsourcing the decision on who can see their website to a large company (Cloudflare). If the company decides that a certain user should not see the website, the user will not see the website, and no one will know about it, and the user will have no recourse. That is not the open web that I would like to see. A second side effect of a knee-jerk reaction to bots…

How do you monetise bot traffic?

Re: A year of fighting scrapers on my 1.5 million-page website

#117
post #101

The worrying thing here is that so many people accepted outsourcing the decision on who can see their website to a large company (Cloudflare). If the company decides that a certain user should not see the website, the user will not see the website, and no one will know about it, and the user will have no recourse. That is not the open web that I would like to see. A second side effect of a knee-jerk reaction to bots…

> If I run a local script or an LLM that needs to fetch a web page from your site and you block that script or LLM as being a bot, you hurt me, the user.

No; in this case you are not a user, you are a bot user.

Re: A year of fighting scrapers on my 1.5 million-page website

#118
Can someone help me understand the underlying motivation behind this?

It makes sense that some crawlers, in the style of Google, would want to index the entire internet. But what is the point of the same crawler re-fetching a page they already fetched an hour ago? Or possibly all this traffic is just independent entities, each trying to cache the internet? The scale of bot traffic makes this seem unlikely.

What's the motivation behind the same entity re-fetching a page it just fetched less than an hour ago?

Re: A year of fighting scrapers on my 1.5 million-page website

#119
post #66
post #57

> My normal bill for running this whole site is around $90 a month. During one bad spike month, it jumped about 500%. This is D1 - which has very surprising costs. you may just want to drop D1 and move to a static site. There’s no reason your site should cost this much.

There's a lot of people who host their site at extremely expensive places and then do everything they can to minimise unneeded traffic - instead of just moving to a cheaper host. Vercel is another popular extremely expensive host.

[dead]
Post reply on HN