Earlier quoted context omitted.
What if I want my content freely available to humans, and not to bots? Why is that such an insane, unworkable ask? All I want is a copyleft protection that specifically allows humans to access my work to their heart's content, but disallows AI use of it in any form. Is that truly so unreasonable?
It's not unreasonable to ask but I think it probably is unreasonable to expect a strictly technical solution. It feels like we're in the realm of politics, policy, and law.
The web does not need gatekeepers: Cloudflare’s new “signed agents” pitch
491–500 of 520 posts
Re: The web does not need gatekeepers: Cloudflare’s new “signed agents” pitch
#492Re: The web does not need gatekeepers: Cloudflare’s new “signed agents” pitch
#493Re: The web does not need gatekeepers: Cloudflare’s new “signed agents” pitch
#494While I concur with the effective tech, I don't think this is something that's a net win for society. Just because you can, doesn't mean you should and I don't feel any one entity (private or public) should be an arbiter on these matters. This is something that can, and should, be negotiated at the "last virtual mile".
>Just because you can, doesn't mean you should and I don't feel any one entity (private or public) should be an arbiter on these matters. What do you mean by private? Should I not be allowed to block AI agents on my sites using Cloudflare?
That said, if you're using Cloudflare Workers, where the endpoint itself is hosted by Cloudflare, that would be ethical.
Re: The web does not need gatekeepers: Cloudflare’s new “signed agents” pitch
#495Earlier quoted context omitted.
>Just because you can, doesn't mean you should and I don't feel any one entity (private or public) should be an arbiter on these matters. What do you mean by private? Should I not be allowed to block AI agents on my sites using Cloudflare?
More precisely, Cloudflare should not be able to offer this service. AI agents should be blocked at the hosted API endpoint ("last mile"), not at a CDN or other sort of intermediary. That said, if you're using Cloudflare Workers, where the endpoint itself is hosted by Cloudflare, that would be ethical.
Re: The web does not need gatekeepers: Cloudflare’s new “signed agents” pitch
#496Re: The web does not need gatekeepers: Cloudflare’s new “signed agents” pitch
#497I don't agree that bot identity is a problem or something that we are better off with if it is solved. I'd rather have a web where adversarial interoperability is possible than one where service operators have a say what tools you can use to access their websites.
The main problem with we are see today is bad actors that
a) completely ignore and/or side step copyright an licensing to use work of others for their own benefit without contributing anything back
b) send an unreasonable number of request
Both will not be solved with identifying bots, that will at best get rid of the small players and give Google, Meta, etc. even more power. The unsustainable parasitic theft of open content is something that needs to be dealt with legally and nothing else will solve it. DRM never works. If enough people block Gemini, Google will just feed it with Google bot crawls and no one can afford to block that. And then they will sell the data to other players. Or someone will make a browser extension to do the same.
The second issue also should be solved via legislation and enforcement thereof. It can also be solved by disconnecting and/or throttling abusive networks wholesale - whole countries if need be. You know, like we have been handling abusive network participants forever. Trying to detect "bots" is a fools errand that will only get rid of the laziest crawlers. You cannot win the bot blocking game when the bots can afford to spend more resources per request than real users are willing to.
So yes, the web must remain open. But to do that we must not have bot identity checks at all - whether that's managed by a single company that has inserted it as a gatekeeper or not.
Re: The web does not need gatekeepers: Cloudflare’s new “signed agents” pitch
#498Earlier quoted context omitted.
> Everyone loves the dream of a free for all and open web... But the reality is how can someone small protect their blog or content from AI training bots? Aren't these statements entirely in conflict? You either have a free for all open web or you don't. Blocking AI training bots is not free and open for all.
No, that is not true. It is only true if you just equate "AI training bots" with "people" on some kind of nominal basis without considering how they operate in practice. It is like saying "If your grocery store is open to the public, why is it not open to this herd of rhinoceroses?" Well, the reason is because rhinoceroses are simply not going to stroll up and down the aisles and head to the checkout line quietly wit…
Meanwhile if you are concerned with the parasitic nature of AI companies then no technical measure will solve that. As you have already noted, they can just buy your data from someone else who you can't afford to block - Google, users with a browser extension that records everything, bots that are ahead of you in the game of cat and mouse, etc.
Re: The web does not need gatekeepers: Cloudflare’s new “signed agents” pitch
#499Earlier quoted context omitted.
ISPs are supposed to disconnect abusive customers. The correct thing to do is probably contact the ISP. Don't complain about scraping, complain about the DDOS (which is the actual problem and I'm increasingly beginning to believe the intent.)
Sure, let me just contact that one ISP located in Russia or India, I am sure they will care a lot about my self-hosted blog
Imagine if Chinese or Russian criminal gangs started sending mail bombs to the US/EU and our solution would be to require all senders, including domestic ones, to prove their identity in order to have their parcels delivered. Completely absurd, but somehow with the Internet everyone jumps to that instead of more reasonable solutions.
Re: The web does not need gatekeepers: Cloudflare’s new “signed agents” pitch
#500Earlier quoted context omitted.
You have a problem with badly behaved scrapers, not AI. I can't disagree with being against badly behaved scrapers. But this is neither a new problem or an interesting one from the idea of making information freely available to everyone, even rhinoceroses, assuming they are well behaved. Blocking bad actors is not the same thing as blocking AI.
But many people feel that the very act of incorporating your copyrighted words into their for-profit training set is itself the bad behavior. It's not about rate-limiting scrapers, it's letting them in the door in the first place.
All you end up doing is participating in the enshittification of the web for the rest of us.