Live data from Hacker News

Stay discoverable in search while disallowing AI training

blog.cloudflare.com

11–20 of 59 posts

Re: Stay discoverable in search while disallowing AI training

#12
post #7
post #6

what weirds me out is the analytics theres no way my index.html page with nothing is getting 10000 hits a day wtf?

7 per minute. I use my high school’s website to test Internet connectivity bc the domain is short and they don’t do a TLS redirect (making it easy to detect WiFi portals).

neverssl.com

Re: Stay discoverable in search while disallowing AI training

#13
post #3

A bit too late honestly (?). With so many people who have shifted over to reading AI summaries as a primary search response, those with AI-enabled sites will win by attrition. There is no going back from this. And the internet is a relatively new phenomenon. Recklessly, blindly applying ads to pages in hopes of generating revenue is a very silly thing to do. Technology with ad blockers and now AI summaries has taken…

All this attitude does is tear down the only viable income source for independent publishers and demonizes them for trying to make money, while everyone let's huge corporations off the hook for it because "well that's just what they do"

Re: Stay discoverable in search while disallowing AI training

#14
I maintain a cloud IP ranges database, and I'm going to test this out.

I have my doubts, though. A formal title like "Accountable" (capitalized) sounds deliberate, but I can't help imagining the renewal email:

"Hey, want to renew your Accountable™ license? Just pinky promise again that you use your IPs for what you say you do."

Re: Stay discoverable in search while disallowing AI training

#15

I wonder if protocols like Web Bot Auth [1] will see wider adoption. At least as a supported mechanism for those bots which identify themselves. The rest probably still have to be treated with Anubis. In my free time I've recently been experimenting with a Web Bot Auth implementation as an Envoy dynamic module [2] to have a way to define some additional policies for the traffic from bots. [1] https://datatracker.ietf…

I'm playing around with it for my MCP hiring protocol ojcp[1] and it seems to work very well for signing attestations at the header level.

[1] https://github.com/ojcp-org/ojcp

Re: Stay discoverable in search while disallowing AI training

#16
post #3

A bit too late honestly (?). With so many people who have shifted over to reading AI summaries as a primary search response, those with AI-enabled sites will win by attrition. There is no going back from this. And the internet is a relatively new phenomenon. Recklessly, blindly applying ads to pages in hopes of generating revenue is a very silly thing to do. Technology with ad blockers and now AI summaries has taken…

All this attitude does is tear down the only viable income source for independent publishers and demonizes them for trying to make money, while everyone let's huge corporations off the hook for it because "well that's just what they do"

Advertising in the way it’s done is demonic in and of itself. I don’t care - find a better business model.

Re: Stay discoverable in search while disallowing AI training

#17

I wonder if protocols like Web Bot Auth [1] will see wider adoption. At least as a supported mechanism for those bots which identify themselves. The rest probably still have to be treated with Anubis. In my free time I've recently been experimenting with a Web Bot Auth implementation as an Envoy dynamic module [2] to have a way to define some additional policies for the traffic from bots. [1] https://datatracker.ietf…

I'm playing around with it for my MCP hiring protocol ojcp[1] and it seems to work very well for signing attestations at the header level. [1] https://github.com/ojcp-org/ojcp

Thanks for sharing, it is interesting use case

Re: Stay discoverable in search while disallowing AI training

#20
"Cloudflare classifies bots by behavior, and a single bot can exhibit more than one behavior."

Is that really true

CF classifies anyone not using a popular browser with Javascript enabled as a "bot"

CF fingerprints www users

As an example, look at CF's Permissions-Policy HTTP response header on a site with CF "bot protection", i.e., the "checking your browser" CAPTCHA nonsense (challenges.cloudflare.com). Then look at IA's Permissions-Policy response header. One CDN is advertiser-focused, the other is user-focused

IA = Internet Archive

Post reply on HN