Live data from Hacker News

The web does not need gatekeepers: Cloudflare’s new “signed agents” pitch

positiveblue.substack.com

491–500 of 520 posts

Re: The web does not need gatekeepers: Cloudflare’s new “signed agents” pitch

#491

Earlier quoted context omitted.

What if I want my content freely available to humans, and not to bots? Why is that such an insane, unworkable ask? All I want is a copyleft protection that specifically allows humans to access my work to their heart's content, but disallows AI use of it in any form. Is that truly so unreasonable?

It's not unreasonable to ask but I think it probably is unreasonable to expect a strictly technical solution. It feels like we're in the realm of politics, policy, and law.

Oh, sure. I absolutely want a legal solution, not a technical one.

Re: The web does not need gatekeepers: Cloudflare’s new “signed agents” pitch

#492
post #485

Earlier quoted context omitted.

6 months from now when you've internalized this entire thread are you even going to remember where you got it from?

Why are you shifting the discussion by adding two new variables (time/memory)?

Because that's how one interacts with AI.

Re: The web does not need gatekeepers: Cloudflare’s new “signed agents” pitch

#494

While I concur with the effective tech, I don't think this is something that's a net win for society. Just because you can, doesn't mean you should and I don't feel any one entity (private or public) should be an arbiter on these matters. This is something that can, and should, be negotiated at the "last virtual mile".

>Just because you can, doesn't mean you should and I don't feel any one entity (private or public) should be an arbiter on these matters. What do you mean by private? Should I not be allowed to block AI agents on my sites using Cloudflare?

More precisely, Cloudflare should not be able to offer this service. AI agents should be blocked at the hosted API endpoint ("last mile"), not at a CDN or other sort of intermediary.

That said, if you're using Cloudflare Workers, where the endpoint itself is hosted by Cloudflare, that would be ethical.

Re: The web does not need gatekeepers: Cloudflare’s new “signed agents” pitch

#495

Earlier quoted context omitted.

>Just because you can, doesn't mean you should and I don't feel any one entity (private or public) should be an arbiter on these matters. What do you mean by private? Should I not be allowed to block AI agents on my sites using Cloudflare?

More precisely, Cloudflare should not be able to offer this service. AI agents should be blocked at the hosted API endpoint ("last mile"), not at a CDN or other sort of intermediary. That said, if you're using Cloudflare Workers, where the endpoint itself is hosted by Cloudflare, that would be ethical.

If the CDN is implementing the explicitly-configured preferences of the specific site, what difference does it make if the blocking happens at the CDN vs. the site's origin server?

Re: The web does not need gatekeepers: Cloudflare’s new “signed agents” pitch

#497
> Yes, identity for agents is a real problem.

I don't agree that bot identity is a problem or something that we are better off with if it is solved. I'd rather have a web where adversarial interoperability is possible than one where service operators have a say what tools you can use to access their websites.

The main problem with we are see today is bad actors that

a) completely ignore and/or side step copyright an licensing to use work of others for their own benefit without contributing anything back

b) send an unreasonable number of request

Both will not be solved with identifying bots, that will at best get rid of the small players and give Google, Meta, etc. even more power. The unsustainable parasitic theft of open content is something that needs to be dealt with legally and nothing else will solve it. DRM never works. If enough people block Gemini, Google will just feed it with Google bot crawls and no one can afford to block that. And then they will sell the data to other players. Or someone will make a browser extension to do the same.

The second issue also should be solved via legislation and enforcement thereof. It can also be solved by disconnecting and/or throttling abusive networks wholesale - whole countries if need be. You know, like we have been handling abusive network participants forever. Trying to detect "bots" is a fools errand that will only get rid of the laziest crawlers. You cannot win the bot blocking game when the bots can afford to spend more resources per request than real users are willing to.

So yes, the web must remain open. But to do that we must not have bot identity checks at all - whether that's managed by a single company that has inserted it as a gatekeeper or not.

Re: The web does not need gatekeepers: Cloudflare’s new “signed agents” pitch

#498

Earlier quoted context omitted.

> Everyone loves the dream of a free for all and open web... But the reality is how can someone small protect their blog or content from AI training bots? Aren't these statements entirely in conflict? You either have a free for all open web or you don't. Blocking AI training bots is not free and open for all.

No, that is not true. It is only true if you just equate "AI training bots" with "people" on some kind of nominal basis without considering how they operate in practice. It is like saying "If your grocery store is open to the public, why is it not open to this herd of rhinoceroses?" Well, the reason is because rhinoceroses are simply not going to stroll up and down the aisles and head to the checkout line quietly wit…

But we should also not throw out the baby with the bathwater. All these attempts at blocking AI bots also block other kinds of crawlers as well as real users with niche browsers.

Meanwhile if you are concerned with the parasitic nature of AI companies then no technical measure will solve that. As you have already noted, they can just buy your data from someone else who you can't afford to block - Google, users with a browser extension that records everything, bots that are ahead of you in the game of cat and mouse, etc.

Re: The web does not need gatekeepers: Cloudflare’s new “signed agents” pitch

#499

Earlier quoted context omitted.

ISPs are supposed to disconnect abusive customers. The correct thing to do is probably contact the ISP. Don't complain about scraping, complain about the DDOS (which is the actual problem and I'm increasingly beginning to believe the intent.)

Sure, let me just contact that one ISP located in Russia or India, I am sure they will care a lot about my self-hosted blog

Except that's exactly what you should do. And if they refuse to cooperate you contact the network operators between them and yourself.

Imagine if Chinese or Russian criminal gangs started sending mail bombs to the US/EU and our solution would be to require all senders, including domestic ones, to prove their identity in order to have their parcels delivered. Completely absurd, but somehow with the Internet everyone jumps to that instead of more reasonable solutions.

Re: The web does not need gatekeepers: Cloudflare’s new “signed agents” pitch

#500
post #254

Earlier quoted context omitted.

You have a problem with badly behaved scrapers, not AI. I can't disagree with being against badly behaved scrapers. But this is neither a new problem or an interesting one from the idea of making information freely available to everyone, even rhinoceroses, assuming they are well behaved. Blocking bad actors is not the same thing as blocking AI.

But many people feel that the very act of incorporating your copyrighted words into their for-profit training set is itself the bad behavior. It's not about rate-limiting scrapers, it's letting them in the door in the first place.

But that is not something you can protect against with technical means. At beast you can block the little fish and give even more power to the mega corporations who will always have a way to get to the data - either by operating crawlers you cannot afford to block, incentivizing users to run their browsers and/or extensions that collect the data and/or buying the data from someone who does.

All you end up doing is participating in the enshittification of the web for the rest of us.

Post reply on HN