Live data from Hacker News

Cloudflare Introduces Default Blocking of A.I. Data Scrapers

nytimes.com

301–310 of 342 posts

Re: Cloudflare Introduces Default Blocking of A.I. Data Scrapers

#301
post #120

Earlier quoted context omitted.

How is Cloudflare a parasite? I can use Cloudflare, and get their AI protection, for free. I have dozens of domains I have used with Cloudflare at one point and I haven't paid them a dime.

Download Brave. Turn on Tor and browse for a week. Now you know what “undesirables” feel like, where “undesirables” can be from a poor country, a bad IP block, outdated browsers, etc. It sucks.

Why download brave and the use Tor.

Just use the Tor browser

Re: Cloudflare Introduces Default Blocking of A.I. Data Scrapers

#302
post #243
post #159

Earlier quoted context omitted.

We need a reasonable alternative to some of what Cloudflare does that can be easily installed as a package on Linux distributions without any of the following to install it. * curl | bash * Docker * Anything that smacks of cryptocurrency or other scams Just a standard repo for Debian and RHEL derived distros. Fully open source so everyone can use it. (apt/dnf install no-bad-actors) Until that exists, using Cloudflare…

I make this: https://anubis.techaro.lol . I have yet to add the SQL injection or IP list layers, but I can add that to the roadmap.

Primary reason people use cloudflare is to hide the ip address of their own server. So they are less likely to be hacked.

Most people are not worried about DDos as their is no reason for any one to DDos them.

Until other services start offering the same, Cloudflare remains default.

Re: Cloudflare Introduces Default Blocking of A.I. Data Scrapers

#303
post #128

I turned this on and it adjusts the robots.txt automatically; not sure what else it is doing. # NOTICE: The collection of content and other data on this # site through automated means, including any device, tool, # or process designed to data mine or scrape content, is # prohibited except (1) for the purpose of search engine indexing or # artificial intelligence retrieval augmented generation or (2) with express # wr…

what actually are the consequences of ignoring robots.txt (apart from DDOS)? have any of these cases ended up in court at all?

BBC recently served a cease and desist on perplexity to stop, and delete all existing.

https://www.bbc.co.uk/news/articles/cy7ndgylzzmo

So an ai company can just be naughty till asked to stop, and then exclud that one company that has the financial resources to go legal.

Re: Cloudflare Introduces Default Blocking of A.I. Data Scrapers

#304
post #190
post #115

Earlier quoted context omitted.

For my silly hobby sites I just return status 444 close the connection for anything that has case-insentive "bot" in the UA requesting anything other than robots.txt, humans.txt, favicon.ico, etc... This would also drop search engines but I blackhole route most of their CIDR blocks. I'm probably the only one here that would do this.

How does a bot scraping your silly hobby sites for any purpose harm or negatively affect you in any way?

Only if they push me over my bandwidth limits but they can't do that if I just drop them on the floor.

Re: Cloudflare Introduces Default Blocking of A.I. Data Scrapers

#305
post #83

Earlier quoted context omitted.

They arent doing anything. They are attempting to insert themselves into the middle of a marketplace (that doesnt exist and never will) where scrapers pay for IP. They think theyre going to profit off the bots, not protect your site. Dont fall for their scam.

What do you mean they are trying to insert themselves? If I have a website that I host with cloudflare, I (as the rightful website owner) has inserted Cloudflare in between. It isnt CF going around saying, that's a nice website you have there. I'm gonna put myself in between.

[deleted]

Re: Cloudflare Introduces Default Blocking of A.I. Data Scrapers

#306

Earlier quoted context omitted.

Auth? Because whatever Cloudflare is doing isn't going to stop anyone serious about scraping data.

Let’s say I’m talking about content that I don’t want behind an auth wall. Is your position simply that all such sites should abandon any efforts to not have the content used for LLM training?

CF will stop bots that respect your robots.txt, and try and stop ones that don't. If your concern is just that you don't want your content used to train an LLM, this will stop the honest companies.

If you are concerned about load on your site because the crawlers are hammering your site, the ones that respect robots.txt should be respecting your crawl delay too. CF will be able to block the dumb ones that ignore your robots.txt and hammer you with no real strategy.

But serious scrapers will have rotating residential IPs and be loading your site from real browsers, they'll take effort to appear as actual users. Sites like Ticket Master have an endless arms race against these. Some Chinese LLM company will get your data if it's public lol.

Re: Cloudflare Introduces Default Blocking of A.I. Data Scrapers

#307
post #201

Earlier quoted context omitted.

Not sure if you're joking, but if you're not: Congratulations on using a very "normal/safe" OS/browser/IP. I get captchas daily, without using any VPN and on several different IPs (work, home, mobile). The only crime I can think of is that I'm using Firefox instead of Chrome.

Since a few days ago, I've been getting Captchas hourly or more. It's probably because I use Firefox on Linux with an ad blocker. For my part, I've ensured we don't use Cloudflare at work.

Using Linux is rare among the general public, but very normal among the kind of person who may find themselves working at Cloudflare or at a potential cloudflare partner/customer.

I don't really buy the argument that they're pushing more captchas to you just because of using Firefox on Linux with an ad blocker.

Re: Cloudflare Introduces Default Blocking of A.I. Data Scrapers

#308
post #301

Earlier quoted context omitted.

Download Brave. Turn on Tor and browse for a week. Now you know what “undesirables” feel like, where “undesirables” can be from a poor country, a bad IP block, outdated browsers, etc. It sucks.

Why download brave and the use Tor. Just use the Tor browser

Some large percentage of people fail when directed to the Tor browser; I don't know why.

Re: Cloudflare Introduces Default Blocking of A.I. Data Scrapers

#310

Earlier quoted context omitted.

Tbh, that content I'm mostly fine with. My only real issue is that people are making trillions off the free labor of people like you and me, giving less time to create that OSS and blogs. But this isn't new to AI, it is just scaled. What I do care about is the theft of my identity. A person may learn from the words I write but that person doesn't end up mimicking the way I write. They are still uniquely themselves. I…

> OSS > people are making trillions off the free labor of people like you and me I read "No Discrimination Against Fields of Endeavor" to also include LLMs and especially the cases that we most deeply disagree with. Either we believe in the principles of OSS or we do not. If you do not like the idea of your intellectual property being used for commercial purposes then this model is definitely not for you. There is no…

I think this may be too much of a "literal" interpretation of OSS without really considering the social contract many OSS supporters believe in, wherein users of OSS will act in good faith and might eventually reciprocate for the benefits they're getting, e.g. the way companies have slowly accepted paying their own employees to contribute to projects openly, releasing their own open source code, respecting the spirit of OSS licenses, sponsoring the developers of the thing they use, etc.

I think it's entirely fair that even staunch supporters of OSS get turned off when AI companies scrape their work to ingest into a black box regurgitator and then turn around and tell the world how their AI will make trillions of dollars by taking away the jobs of those obsolete OSS developers, showing no intention of ever giving back to the community.

Post reply on HN