Live data from Hacker News

Cloudflare Introduces Default Blocking of A.I. Data Scrapers

nytimes.com

171–180 of 342 posts

Re: Cloudflare Introduces Default Blocking of A.I. Data Scrapers

#171
post #13

> If A.I. companies freely use data from various websites without permission or payment, people will be discouraged from creating new digital content I don't see a way out of this happening. AI fundamentally discourages other forms of digital interaction as it grows. Its mechanism of growing is killing other kinds of digital content. It will eventually kill the web, which is, ironically, its main source of food.

[flagged]

nothing's perfect but it is still better than the other options

Re: Cloudflare Introduces Default Blocking of A.I. Data Scrapers

#172
post #164
post #147

Earlier quoted context omitted.

Right - do I want them getting some info from me, or do I want my IP address exposed? Besides CloudFront, which still costs money, what other option is there for semi-privacy and caching for free?

As the old addage goes: If you're not paying for it, you're the product. Lots of nuance, but generally: pay for things you use. Servers, engineers, and research and development are not free, so someone has to pay.

Lots of services don't even let me pay if I wanted to, so I am forced to be the product. (Donating typically does not un-productify myself).

Or I pay and am still the product. Just with less in-my-face ads.

Re: Cloudflare Introduces Default Blocking of A.I. Data Scrapers

#173
post #137

Earlier quoted context omitted.

They put themselves as a middle man for almost the whole Internet, collect huge usage data about everyone and block anybody who doesn't use mainstream tools: https://news.ycombinator.com/item?id=42953508 https://news.ycombinator.com/item?id=13718752 https://news.ycombinator.com/item?id=23897705 https://news.ycombinator.com/item?id=41864632 https://news.ycombinator.com/item?id=42577076

valid. ...but OTOH it's their customers who want all of that and pay to get that, because the alternative is worse. rock and a hard place.

I want to know if there is a way to design an alternative that isn't controlled by a single entity which allows gatekeeping.

Re: Cloudflare Introduces Default Blocking of A.I. Data Scrapers

#174

Few people realise that virtually everything we do online has, until this point, been free training to make OpenAI, Anthropic, etc. richer while cutting humans--the ones who produced the value--out of the loop. It might be too little, too late, at this juncture, and this particular solution doesn't seem too innovative. However, it is directionally 100% correct, and let's hope for massively more innovation in defendin…

[flagged]

Re: Cloudflare Introduces Default Blocking of A.I. Data Scrapers

#175
post #147
post #137

Earlier quoted context omitted.

valid. ...but OTOH it's their customers who want all of that and pay to get that, because the alternative is worse. rock and a hard place.

Right - do I want them getting some info from me, or do I want my IP address exposed? Besides CloudFront, which still costs money, what other option is there for semi-privacy and caching for free?

Cloud front is pretty much free for your first TB. Fastly has a free plan.

Though why should it be for free?

Re: Cloudflare Introduces Default Blocking of A.I. Data Scrapers

#176

Earlier quoted context omitted.

Tbh, that content I'm mostly fine with. My only real issue is that people are making trillions off the free labor of people like you and me, giving less time to create that OSS and blogs. But this isn't new to AI, it is just scaled. What I do care about is the theft of my identity. A person may learn from the words I write but that person doesn't end up mimicking the way I write. They are still uniquely themselves. I…

> OSS > people are making trillions off the free labor of people like you and me I read "No Discrimination Against Fields of Endeavor" to also include LLMs and especially the cases that we most deeply disagree with. Either we believe in the principles of OSS or we do not. If you do not like the idea of your intellectual property being used for commercial purposes then this model is definitely not for you. There is no…

  > Either we believe in the principles of OSS or we do not.
What about respecting licenses?

Seriously, don't lick the boot. We can recognize that there's complexity here. Trivializing everything only helps the abusers.

Giving credit where credit is due is not too much to ask. Other people making money off my work can be good[0]. Taking credit for it is insulting

[0] If you're not making much, who cares. But if you're a trillion dollar business you can afford to give a little back. Here's the truth, OSS only works if we get enough money and time to do the work. That's either by having a good work life balance and good pay or enough donations coming in. We've been mostly supported by the former, but that deal seems to be going away

Re: Cloudflare Introduces Default Blocking of A.I. Data Scrapers

#178

Earlier quoted context omitted.

LLM scrapers have dramatically been increasing the cost of hosting various small websites. Without something being done, the data that these scrapers rely on would eventually no longer exist.

I think the correct term is, that unrestricted LLM scrapers have dramatically been increasing the cost of hosting various small websites. Its not a issue when somebody does "ethical" scraping, with for instance, a 250ms delay between requests, and a active cache that checks specific pages (like news article links) to rescrape at 12 or 24h intervals. This type of scraping results in almost no pressure on the websites.…

The biggest issue for me is clearly masquerading their User-Agent strings. Regardless of whether they are slow and respectful crawlers, they should clearly identify themselves, provide a documentation URL and obey robots.txt. Without that, I have to play a frankly tiring game of cat and mouse, wasting my time and the time of my users (they have to put up with some form of captcha or PoW thing).

I've been an active lurker in the self-hosting community and I'm definitely not alone. Nearly everyone hosting public facing websites, particularly those whose form is rather juicy for LLMs, have been facing these issues. It costs more time and money to deal with this, when applying a simple User-Agent block would be much cheaper and trivial to do and maintain.

sigh

Re: Cloudflare Introduces Default Blocking of A.I. Data Scrapers

#179

Earlier quoted context omitted.

Which companies? OpenAI, Anthropic, Google? No, their bots are pretty well behaved. The smaller AI companies deploying bots that don't respect any reasonable rate limits and are scraping the same static pages thousands of times an hour? Yup

Anecdote, but at least for tiny little server hosting single public repository, none of these companies had 'well behaved' bots. It may be possible that they learned to behave better but I wouldn't know since my only possible recourse was to blacklist them all AND take the repo private.

Those are the small companies spoofing their user agent as the big companies to dodge countermeasures.

Re: Cloudflare Introduces Default Blocking of A.I. Data Scrapers

#180
post #175
post #147

Earlier quoted context omitted.

Right - do I want them getting some info from me, or do I want my IP address exposed? Besides CloudFront, which still costs money, what other option is there for semi-privacy and caching for free?

Cloud front is pretty much free for your first TB. Fastly has a free plan. Though why should it be for free?

Multiple people have brought that up. I pay for everything else, why not one more.

Although bunny.net won't take ANY of my credit or debit cards

Post reply on HN