Live data from Hacker News

Stay discoverable in search while disallowing AI training

blog.cloudflare.com

51–57 of 57 posts

Re: Stay discoverable in search while disallowing AI training

#51

I feel like most of these schemes to categorise data as "public but not really" are ultimately doomed to failure. Even if you could trust every AI company in the world to respect these terms, is there anything stopping someone else indexing the data and selling them the information? I know there's copyright law but they're apparently ignoring that anyway. Ultimately this reminds me of those really early social media…

It's also a way to literally advertise to AI companies that you have some data worth plundering in an increasingly dead and sloppy internet that has diminishing returns for training.

Re: Stay discoverable in search while disallowing AI training

#52
post #24

Earlier quoted context omitted.

As far as I can tell, after months of fighting being DDoSed by Anthropic and OpenAI across 50+ sites - Cloudflare also allows what it considers "good bots" through all of your bot blocking rules, with no option to turn this off unless you pay them money.

How did you dertmine it was those two companies? Also did you disallow them in robots.txt?

No, I definitely spent months fighting off bots across multiple hosts and never set up a robots.txt file anywhere nor did I look at the analytics dashboard on cloudflare.

Re: Stay discoverable in search while disallowing AI training

#54

"Accountable" is just a fancy word for "pinky promise, but with a label." Nothing stops the data from ending up in a training run once it's already been fetched.

> "pinky promise, but with a label."

and that label is "pending litigation"

Re: Stay discoverable in search while disallowing AI training

#55

"Cloudflare classifies bots by behavior, and a single bot can exhibit more than one behavior." Is that really true CF classifies anyone not using a popular browser with Javascript enabled as a "bot" CF fingerprints www users As an example, look at CF's Permissions-Policy HTTP response header on a site with CF "bot protection", i.e., the "checking your browser" CAPTCHA nonsense (challenges.cloudflare.com). Then look a…

[dead]

Re: Stay discoverable in search while disallowing AI training

#56

I feel like most of these schemes to categorise data as "public but not really" are ultimately doomed to failure. Even if you could trust every AI company in the world to respect these terms, is there anything stopping someone else indexing the data and selling them the information? I know there's copyright law but they're apparently ignoring that anyway. Ultimately this reminds me of those really early social media…

> If you don't want your content to end up in some database/archive don't publish it for the whole world to see.

This principle somewhat reminds me of the line that "If you're not paying for the product, you are the product", and it seems to me similarly misleading - my data gets harvested and sold by companies with which I have non-paying relationships and by companies I have to pay for things (I am made the product in both cases). As you note - the AI companies are ignoring copyright law and pirating everything that seems useful to them regardless of whether it was published for free access.

The potential externalities here are troubling.

https://vbuckenham.com/blog/how-to-find-things-online/

Re: Stay discoverable in search while disallowing AI training

#57
post #52

Earlier quoted context omitted.

How did you dertmine it was those two companies? Also did you disallow them in robots.txt?

No, I definitely spent months fighting off bots across multiple hosts and never set up a robots.txt file anywhere nor did I look at the analytics dashboard on cloudflare.

I'm asking sincerely, how do you get to the conclusion that it's those two?
Post reply on HN