Live data from Hacker News

Cloudflare Introduces Default Blocking of A.I. Data Scrapers

nytimes.com

161–170 of 342 posts

Re: Cloudflare Introduces Default Blocking of A.I. Data Scrapers

#161

Earlier quoted context omitted.

Tbh, that content I'm mostly fine with. My only real issue is that people are making trillions off the free labor of people like you and me, giving less time to create that OSS and blogs. But this isn't new to AI, it is just scaled. What I do care about is the theft of my identity. A person may learn from the words I write but that person doesn't end up mimicking the way I write. They are still uniquely themselves. I…

> OSS > people are making trillions off the free labor of people like you and me I read "No Discrimination Against Fields of Endeavor" to also include LLMs and especially the cases that we most deeply disagree with. Either we believe in the principles of OSS or we do not. If you do not like the idea of your intellectual property being used for commercial purposes then this model is definitely not for you. There is no…

Open source software typically has a license. People not following the license isn’t tolerated.

This is what AI scrapers are doing. They’re taking your code, your artwork and your writing without any consideration for the license.

Re: Cloudflare Introduces Default Blocking of A.I. Data Scrapers

#162
post #130

My data served by Cloudflare has increased to 100gb /month compared to <20gb like 2 years ago, and they're all fairly static hobby sites. Actual people traffic is down by like half in the same time frame, so I imagine a lot of this is probably cost savings for Cloudflare to reduce resource usage.

Makes total sense, bandwidth on this scale is expensive.

Re: Cloudflare Introduces Default Blocking of A.I. Data Scrapers

#163

Earlier quoted context omitted.

Auth? Because whatever Cloudflare is doing isn't going to stop anyone serious about scraping data.

Let’s say I’m talking about content that I don’t want behind an auth wall. Is your position simply that all such sites should abandon any efforts to not have the content used for LLM training?

If you find a solution that’s not auth please let me know.

Re: Cloudflare Introduces Default Blocking of A.I. Data Scrapers

#164
post #147
post #137

Earlier quoted context omitted.

valid. ...but OTOH it's their customers who want all of that and pay to get that, because the alternative is worse. rock and a hard place.

Right - do I want them getting some info from me, or do I want my IP address exposed? Besides CloudFront, which still costs money, what other option is there for semi-privacy and caching for free?

As the old addage goes: If you're not paying for it, you're the product.

Lots of nuance, but generally: pay for things you use. Servers, engineers, and research and development are not free, so someone has to pay.

Re: Cloudflare Introduces Default Blocking of A.I. Data Scrapers

#165

Few people realise that virtually everything we do online has, until this point, been free training to make OpenAI, Anthropic, etc. richer while cutting humans--the ones who produced the value--out of the loop. It might be too little, too late, at this juncture, and this particular solution doesn't seem too innovative. However, it is directionally 100% correct, and let's hope for massively more innovation in defendin…

That was always the cost of free and open exchange of ideas though. The idea of the internet in the first place was to allow people to communicate in the open and publish ideas freely. There was never any stipulation that using the published ideas to make money was off limits.

Technology has advanced and now reading the sum total of the freely exchanged ideas has become particularly valuable. But who cares? The internet still exists and is still usable to freely exchange ideas the way it’s always been.

The value that one website provides is a minuscule amount, the value of one individual poster on Reddit is minuscule. Are we asking that each poster on Reddit be paid 1 penny (that’s probably what your posts are worth) for their individual contribution? My websites were used to train these models probably, but the value that each contributed is so small that I wouldn’t even expect a few cents for it.

The person who’s going to profit here is Cloudflare or the owners of Reddit, or any other gatekeeper site that is already profiting from other people’s contributions.

The “parasitism” here just feels like normal competition between giant companies who have special access to information.

Re: Cloudflare Introduces Default Blocking of A.I. Data Scrapers

#166

Earlier quoted context omitted.

Tbh, that content I'm mostly fine with. My only real issue is that people are making trillions off the free labor of people like you and me, giving less time to create that OSS and blogs. But this isn't new to AI, it is just scaled. What I do care about is the theft of my identity. A person may learn from the words I write but that person doesn't end up mimicking the way I write. They are still uniquely themselves. I…

> OSS > people are making trillions off the free labor of people like you and me I read "No Discrimination Against Fields of Endeavor" to also include LLMs and especially the cases that we most deeply disagree with. Either we believe in the principles of OSS or we do not. If you do not like the idea of your intellectual property being used for commercial purposes then this model is definitely not for you. There is no…

> Either we believe in the principles of OSS or we do not. If you do not like the idea of your intellectual property being used for commercial purposes then this model is definitely not for you.

I've been writing open source for more than 20 years

I gave away my work for free with one condition: leave my name on it (MIT license)

the AI parasites then strip the attribution out

they are the ones violating the principles of open source

> then perhaps a different licensing and distribution model is what you are after.

I've now stopped producing open source entirely

and I suggest every developer does the same until the legal position is clarified (in our favour)

Re: Cloudflare Introduces Default Blocking of A.I. Data Scrapers

#167
post #158
post #147

Earlier quoted context omitted.

Right - do I want them getting some info from me, or do I want my IP address exposed? Besides CloudFront, which still costs money, what other option is there for semi-privacy and caching for free?

bunny.net has some options

[deleted]

Re: Cloudflare Introduces Default Blocking of A.I. Data Scrapers

#168
post #158
post #147

Earlier quoted context omitted.

Right - do I want them getting some info from me, or do I want my IP address exposed? Besides CloudFront, which still costs money, what other option is there for semi-privacy and caching for free?

bunny.net has some options

I will have to check them out I guess

Re: Cloudflare Introduces Default Blocking of A.I. Data Scrapers

#169
post #156

Earlier quoted context omitted.

Thank you for having this attitude. I have never attempted any blogging because I always figured no one is actually going to read it. With LLMs, however, I know they will. I actually see this as a motivation to blog, as we are in a position to shape this emerging knowledge base. I don't find it discouraging that others may be profiting off our freely published work, just as I myself have benefited tremendously from o…

This is an interesting take, thanks for sharing. I wonder how someone should adjust their blogging if they believe their primary audience will be LLMs.

SEO -> LLMEO

Re: Cloudflare Introduces Default Blocking of A.I. Data Scrapers

#170
post #112

Earlier quoted context omitted.

It's cloudflare and parasites like them that will make the internet un-free. It's already happening, I'm either blocked or back to 1998 load times be cause of "checking your browser". They are destroying the internet and will make it so only people who do approved things on approved browsers (meaning let advertising companies monetize their online activity) will get real access. Cloudflare isn't solving a problem, th…

LLM scrapers have dramatically been increasing the cost of hosting various small websites. Without something being done, the data that these scrapers rely on would eventually no longer exist.

I think the correct term is, that unrestricted LLM scrapers have dramatically been increasing the cost of hosting various small websites.

Its not a issue when somebody does "ethical" scraping, with for instance, a 250ms delay between requests, and a active cache that checks specific pages (like news article links) to rescrape at 12 or 24h intervals. This type of scraping results in almost no pressure on the websites.

The issue that i have seen, is that the more unscrupulous parties, just let their scrapers go wild, constantly rescraping again and again because the cost of scraping is extreme low. A small VM can easily push 1000's of scraps per second, let alone somebody with more dedicated resources.

Actually building a "ethical" scraper involves more time, as you need to fine tune it per website. Unfortunately, this behavior is going to cost the more ethical scraper a ton, as anti-scraping efforts will increase the cost on our side.

Post reply on HN