Live data from Hacker News

Stay discoverable in search while disallowing AI training

blog.cloudflare.com

41–50 of 59 posts

Re: Stay discoverable in search while disallowing AI training

#42

Earlier quoted context omitted.

All this attitude does is tear down the only viable income source for independent publishers and demonizes them for trying to make money, while everyone let's huge corporations off the hook for it because "well that's just what they do"

Advertising in the way it’s done is demonic in and of itself. I don’t care - find a better business model.

A better business model won’t help. What you need is a better economic and political system. If you “let the market decide” what advertising looks like, you get what we have today, because “let the market decide” is an incoherent claim that’s really shorthand for the rule of a wealthy minority. Advertising is just a convenient way to extract wealth from a society without having to go to all the trouble of satisfying customers.

Re: Stay discoverable in search while disallowing AI training

#44

The irony is that search engines are AI companies now. Telling them 'index me for search but don't train your models' is asking them to split a brain that’s already fully merged.

Exactly. I don’t get this. You may decide to not train but your content will still show up in search.

Re: Stay discoverable in search while disallowing AI training

#45
post #24

Earlier quoted context omitted.

As far as I can tell, after months of fighting being DDoSed by Anthropic and OpenAI across 50+ sites - Cloudflare also allows what it considers "good bots" through all of your bot blocking rules, with no option to turn this off unless you pay them money.

The "good bots" also just happen to be from companies that pay cloudflare a lot of money, I imagine.

Goodness Tokens

Re: Stay discoverable in search while disallowing AI training

#46
post #24

"Cloudflare classifies bots by behavior, and a single bot can exhibit more than one behavior." Is that really true CF classifies anyone not using a popular browser with Javascript enabled as a "bot" CF fingerprints www users As an example, look at CF's Permissions-Policy HTTP response header on a site with CF "bot protection", i.e., the "checking your browser" CAPTCHA nonsense (challenges.cloudflare.com). Then look a…

As far as I can tell, after months of fighting being DDoSed by Anthropic and OpenAI across 50+ sites - Cloudflare also allows what it considers "good bots" through all of your bot blocking rules, with no option to turn this off unless you pay them money.

How did you dertmine it was those two companies? Also did you disallow them in robots.txt?

Re: Stay discoverable in search while disallowing AI training

#47

The irony is that search engines are AI companies now. Telling them 'index me for search but don't train your models' is asking them to split a brain that’s already fully merged.

Even looking for companies which supply services is now far better on AI chats than Google. For me being visible in AI training is going to be more important than search in the next year or two.

If I was a big AI company I'd certainly be tempted to make sure that anyone who excluded themselves from "AI training" also got themselves excluded from AI results.

Re: Stay discoverable in search while disallowing AI training

#48
post #3

A bit too late honestly (?). With so many people who have shifted over to reading AI summaries as a primary search response, those with AI-enabled sites will win by attrition. There is no going back from this. And the internet is a relatively new phenomenon. Recklessly, blindly applying ads to pages in hopes of generating revenue is a very silly thing to do. Technology with ad blockers and now AI summaries has taken…

All this attitude does is tear down the only viable income source for independent publishers and demonizes them for trying to make money, while everyone let's huge corporations off the hook for it because "well that's just what they do"

Papers charged per "user" since times immemorial.

Re: Stay discoverable in search while disallowing AI training

#50
post #9

I didn't know websites could opt out of providing data to Google's AI training. Looks Google added support for this via 'Google-Extended' in robots.txt back in 2023: https://blog.google/innovation-and-ai/products/an-update-on-...

doc: https://developers.google.com/crawling/docs/crawlers-fetcher...
Post reply on HN