Live data from Hacker News

Cloudflare Introduces Default Blocking of A.I. Data Scrapers

nytimes.com

201–210 of 342 posts

Re: Cloudflare Introduces Default Blocking of A.I. Data Scrapers

#201
post #185

Earlier quoted context omitted.

> I have dozens of domains I have used with Cloudflare at one point and I haven't paid them a dime. Maybe you haven't, but your users (primarily those using "suspicious" operating systems and browsers) certainly have – with their time spent solving captchas.

But Cloudflare have removed CAPTCHAs

Not sure if you're joking, but if you're not: Congratulations on using a very "normal/safe" OS/browser/IP.

I get captchas daily, without using any VPN and on several different IPs (work, home, mobile). The only crime I can think of is that I'm using Firefox instead of Chrome.

Re: Cloudflare Introduces Default Blocking of A.I. Data Scrapers

#202

Few people realise that virtually everything we do online has, until this point, been free training to make OpenAI, Anthropic, etc. richer while cutting humans--the ones who produced the value--out of the loop. It might be too little, too late, at this juncture, and this particular solution doesn't seem too innovative. However, it is directionally 100% correct, and let's hope for massively more innovation in defendin…

I think its 100% ok to freely train on public internet data. What is absolutely not ok is to crawl at such an excessive speed that it makes it difficult to host small scale websites. Truly a tragedy of the commons.

Agree. The problem lately is that even if each single scraper is doing so “reasonably,” there are so many individuals and groups doing this that it’s still too onerous for many sites. And of course many are not “reasonable.”

Re: Cloudflare Introduces Default Blocking of A.I. Data Scrapers

#203
post #201

Earlier quoted context omitted.

But Cloudflare have removed CAPTCHAs

Not sure if you're joking, but if you're not: Congratulations on using a very "normal/safe" OS/browser/IP. I get captchas daily, without using any VPN and on several different IPs (work, home, mobile). The only crime I can think of is that I'm using Firefox instead of Chrome.

Since a few days ago, I've been getting Captchas hourly or more.

It's probably because I use Firefox on Linux with an ad blocker.

For my part, I've ensured we don't use Cloudflare at work.

Re: Cloudflare Introduces Default Blocking of A.I. Data Scrapers

#204

Few people realise that virtually everything we do online has, until this point, been free training to make OpenAI, Anthropic, etc. richer while cutting humans--the ones who produced the value--out of the loop. It might be too little, too late, at this juncture, and this particular solution doesn't seem too innovative. However, it is directionally 100% correct, and let's hope for massively more innovation in defendin…

[deleted]

Re: Cloudflare Introduces Default Blocking of A.I. Data Scrapers

#205

Few people realise that virtually everything we do online has, until this point, been free training to make OpenAI, Anthropic, etc. richer while cutting humans--the ones who produced the value--out of the loop. It might be too little, too late, at this juncture, and this particular solution doesn't seem too innovative. However, it is directionally 100% correct, and let's hope for massively more innovation in defendin…

Is it even possible that Cloudfare could manage to block all AI data scrapping? I think this measure is just going to make it harder and more expensive, which will stop AI scrappers from hitting every single page every single day and creating expenses for publishers, but not actually stop their data from ending up in a few datasets.

Re: Cloudflare Introduces Default Blocking of A.I. Data Scrapers

#206
post #110

Earlier quoted context omitted.

I write online (comments here, open source software, blogging, etc) because I have ideas I want to share. Whether it's "I did a thing and here's how" or "we should change policy in this specific way" or "does anyone know how to X" I'm happy for this to go into training models just like I'm happy for it to go into humans reading.

Tbh, that content I'm mostly fine with. My only real issue is that people are making trillions off the free labor of people like you and me, giving less time to create that OSS and blogs. But this isn't new to AI, it is just scaled. What I do care about is the theft of my identity. A person may learn from the words I write but that person doesn't end up mimicking the way I write. They are still uniquely themselves. I…

> A person may learn from the words I write but that person doesn't end up mimicking the way I write.

Oh, I wish I could get AI to mimic the way I write! I'd pay money for it. I often want to type up an email/doc/whatever but don't because of occasional RSI issues. If I could get an AI to type it up for me while still sounding like me - that would be a big boon for my health.

Re: Cloudflare Introduces Default Blocking of A.I. Data Scrapers

#207

Few people realise that virtually everything we do online has, until this point, been free training to make OpenAI, Anthropic, etc. richer while cutting humans--the ones who produced the value--out of the loop. It might be too little, too late, at this juncture, and this particular solution doesn't seem too innovative. However, it is directionally 100% correct, and let's hope for massively more innovation in defendin…

This has been going on even since early social media. I think most of the users actually prefer it.

Re: Cloudflare Introduces Default Blocking of A.I. Data Scrapers

#208

Earlier quoted context omitted.

I wonder… Google scrapes for indexing and for AI, right? I wonder if they will eventually say: ok, you can have me or not, if you don’t want to help train my AI you won’t get my searches either. That’s a tough deal but it is sort of self-consistent.

Very few people seems to be complaining that Google crashes their sites. Google also publish their crawlers IP ranges, but you really don't need to rate-limit Google, they know how to back off and not overload sites.

In theory — in practise I've had to limit Google on two large sites at work. I currently have them limited to 10/s for non-cached requests.

Re: Cloudflare Introduces Default Blocking of A.I. Data Scrapers

#209
post #201

Earlier quoted context omitted.

Not sure if you're joking, but if you're not: Congratulations on using a very "normal/safe" OS/browser/IP. I get captchas daily, without using any VPN and on several different IPs (work, home, mobile). The only crime I can think of is that I'm using Firefox instead of Chrome.

Since a few days ago, I've been getting Captchas hourly or more. It's probably because I use Firefox on Linux with an ad blocker. For my part, I've ensured we don't use Cloudflare at work.

I use firefox on linux with an ad blocker and cloudfare works fine

Re: Cloudflare Introduces Default Blocking of A.I. Data Scrapers

#210
post #32

I've heard lots of people on HN complaining about bot traffic bogging down their websites, and as a website operator myself I'm honestly puzzled. If you're already using Cloudflare, some basic cache configuration should guarantee that most bot traffic hits the cache and doesn't bog down your servers. And even if you don't want to do that, bandwidth and CPU are so cheap these days that it shouldn't make a difference.…

I too am a bit confused / mystified at the strong reaction. But I do expect a lot of badly optimized sites that just want out. I struggle to think of a web related library that has spread faster than Anubis checker. It's everywhere now! https://github.com/TecharoHQ/anubis I'm surprised we don't see more efforts to rate limit. I assume many of these are distributed crawlers, but it feels like there's got to be pools o…

I'm not an expert on website hosting, but after reading some of the blog posts on Anubis, those people were truly at wit's end trying to block AI scrappers with techniques like the ones you imply.
Post reply on HN