I feel like most of these schemes to categorise data as "public but not really" are ultimately doomed to failure. Even if you could trust every AI company in the world to respect these terms, is there anything stopping someone else indexing the data and selling them the information? I know there's copyright law but they're apparently ignoring that anyway. Ultimately this reminds me of those really early social media…
Stay discoverable in search while disallowing AI training
51–57 of 57 posts
Re: Stay discoverable in search while disallowing AI training
#52Earlier quoted context omitted.
As far as I can tell, after months of fighting being DDoSed by Anthropic and OpenAI across 50+ sites - Cloudflare also allows what it considers "good bots" through all of your bot blocking rules, with no option to turn this off unless you pay them money.
How did you dertmine it was those two companies? Also did you disallow them in robots.txt?
Re: Stay discoverable in search while disallowing AI training
#53Re: Stay discoverable in search while disallowing AI training
#54"Accountable" is just a fancy word for "pinky promise, but with a label." Nothing stops the data from ending up in a training run once it's already been fetched.
and that label is "pending litigation"
Re: Stay discoverable in search while disallowing AI training
#55"Cloudflare classifies bots by behavior, and a single bot can exhibit more than one behavior." Is that really true CF classifies anyone not using a popular browser with Javascript enabled as a "bot" CF fingerprints www users As an example, look at CF's Permissions-Policy HTTP response header on a site with CF "bot protection", i.e., the "checking your browser" CAPTCHA nonsense (challenges.cloudflare.com). Then look a…
Re: Stay discoverable in search while disallowing AI training
#56I feel like most of these schemes to categorise data as "public but not really" are ultimately doomed to failure. Even if you could trust every AI company in the world to respect these terms, is there anything stopping someone else indexing the data and selling them the information? I know there's copyright law but they're apparently ignoring that anyway. Ultimately this reminds me of those really early social media…
This principle somewhat reminds me of the line that "If you're not paying for the product, you are the product", and it seems to me similarly misleading - my data gets harvested and sold by companies with which I have non-paying relationships and by companies I have to pay for things (I am made the product in both cases). As you note - the AI companies are ignoring copyright law and pirating everything that seems useful to them regardless of whether it was published for free access.
The potential externalities here are troubling.
Re: Stay discoverable in search while disallowing AI training
#57Earlier quoted context omitted.
How did you dertmine it was those two companies? Also did you disallow them in robots.txt?
No, I definitely spent months fighting off bots across multiple hosts and never set up a robots.txt file anywhere nor did I look at the analytics dashboard on cloudflare.