Live data from Hacker News

A year of fighting scrapers on my 1.5 million-page website

patronview.com

451–460 of 475 posts

Re: A year of fighting scrapers on my 1.5 million-page website

#451

I just checked Cloudflare for SignalBloom ( https://www.signalbloom.ai , which I own and operate). Over the last 72 hours, Claude-searchbot [1] alone fetched ~205,000 pages. Sent exactly 1 referral. There is a lot of free financial data on the site, hoping for real users to benefit from it. It is hard to not feel a little cheated out that Claude gets to claim "Found it!" to its users without me getting no credits or…

Isn't your whole product scraping other sites and summarizing? I also don't see clear sources for your info where it is displayed, so it seems like you do the exact same thing.

If that was it, wouldn't Claude go directly to the canonical source?

Re: A year of fighting scrapers on my 1.5 million-page website

#452

I just checked Cloudflare for SignalBloom ( https://www.signalbloom.ai , which I own and operate). Over the last 72 hours, Claude-searchbot [1] alone fetched ~205,000 pages. Sent exactly 1 referral. There is a lot of free financial data on the site, hoping for real users to benefit from it. It is hard to not feel a little cheated out that Claude gets to claim "Found it!" to its users without me getting no credits or…

FYI, your proof is hosted on a NSFW page. You should have mentioned that.

Apologies, I had no idea. I lookup host image and that site came up. I don't think the site itself is of NSFW nature.

Re: A year of fighting scrapers on my 1.5 million-page website

#453

Anubis[1] is a superb fix for sites not behind Cloudflare/Fastly/Bunny etc. We had millions of bot requests, on a site serving all countries so we couldn't block by country, with fake user-agents so we couldn't block using that. It uses 'proof of work' to detect real browser software. [1] https://anubis.techaro.lol/

It doesn't detect "real browsers". It simply adds some friction and is niche enough that AI scrapers have not bothered to bypass it yet.

> and is niche enough that AI scrapers have not bothered to bypass it yet.

it is not niche. it is state of the art and widely deployed on major websites.

Re: A year of fighting scrapers on my 1.5 million-page website

#454
post #325

Earlier quoted context omitted.

> 10TB/month is usually enough, even with bots. My single static webpage with no updates in 3 years is doing that, which is (one of the reasons) how I end up where that site (and many others in business and personally) is. You're chasing a dream for a world that doesn't exist anymore.

no sorry I don't actually believe your single static webpage is doing 10TB/month. That's about 500 RPS average.

[deleted]

Re: A year of fighting scrapers on my 1.5 million-page website

#455
post #420

Earlier quoted context omitted.

The webmaster is publishing for humans.

Not for browsers then?

When your reach the end of your journey across the sea of semantics and arrive at the shore of a lonely island, you can be king of it. Here, we are examining hosting costs.

Re: A year of fighting scrapers on my 1.5 million-page website

#456
post #455

Earlier quoted context omitted.

Not for browsers then?

When your reach the end of your journey across the sea of semantics and arrive at the shore of a lonely island, you can be king of it. Here, we are examining hosting costs.

I simply don’t get why a human using a browser and a human using a bot using wget to load a page are different for a hoster, from your point of view.

My suspicion is that you did not read the great grand parent post before replying to it and are now engaging in evasive tactics in order not to answer this simple question.

Re: A year of fighting scrapers on my 1.5 million-page website

#457
post #101

The worrying thing here is that so many people accepted outsourcing the decision on who can see their website to a large company (Cloudflare). If the company decides that a certain user should not see the website, the user will not see the website, and no one will know about it, and the user will have no recourse. That is not the open web that I would like to see. A second side effect of a knee-jerk reaction to bots…

> I wanted the information and I have no way of getting at it without investing my own time and effort in going to your website. Substitute the word “website” for book, or training course, or documentary or published paper, or patent or one of hundreds of examples. Now you see the problem.

Huh? I regularly ask codex to not just summarize or extract a crux of published papers, but sometimes to do a topic directed search and give a comparison table. If you aren't doing that, do you see the problem?

Re: A year of fighting scrapers on my 1.5 million-page website

#459
post #6

blocking by geo yeah might work - but what happens when someone is traveling abroad ? they've to use a VPN to access your site ? my take with all the bots - the web is gonna be a bunch of private walled gardens. with most sites set to no index. you will only discover them via referral from someone real.

What happens when someone is from a shit country?

As someone living in Vietnam, I have a number of websites off limits. The one in this thread for one https://www.cam.ac.uk/ is an interesting one that I'm forbidden to see.

VPN is a necessity. I spend some of my day turning it on to see sites blocking Vietnam, and then some of it turning it off because of sites that are blocking the VPN.

Re: A year of fighting scrapers on my 1.5 million-page website

#460

Anubis[1] is a superb fix for sites not behind Cloudflare/Fastly/Bunny etc. We had millions of bot requests, on a site serving all countries so we couldn't block by country, with fake user-agents so we couldn't block using that. It uses 'proof of work' to detect real browser software. [1] https://anubis.techaro.lol/

Because of this wasteful crap the internet is so slow nowadays... Try opening gcc bug tracker on your phone: https://gcc.gnu.org/bugzilla/

Found a better one: https://bugs.winehq.org
Post reply on HN