Live data from Hacker News

The web does not need gatekeepers: Cloudflare’s new “signed agents” pitch

positiveblue.substack.com

401–410 of 520 posts

Re: The web does not need gatekeepers: Cloudflare’s new “signed agents” pitch

#401

Earlier quoted context omitted.

> If you don't want AI bots reading information on the web, you don't actually want a free and open web. This is such a bad faith argument. We want a town center for the whole community to enjoy! What, you don't like those people shooting up drugs over there? But they're enjoying it too, this is what you wanted right? They're not harming you by doing their drugs. Everyone is enjoying it!

If an AI bot is accessing my site the way that regular users are accessing my site -- in other words everyone is using the town center as intended -- what is the problem? Seems to be a lot of conflating of badly coded (intentionally or not) scrapers and AI. That is a problem that predates AI's existence.

So if I buy a DDoS service and DDoS your site, it's ok as long as it accesses it the same way regular people do? In sorry for extreme example, it's obviously not, but that's how I understand your position as written.

We can also consider 10 exploit attempts per second that my site sees.

Re: The web does not need gatekeepers: Cloudflare’s new “signed agents” pitch

#402

While I concur with the effective tech, I don't think this is something that's a net win for society. Just because you can, doesn't mean you should and I don't feel any one entity (private or public) should be an arbiter on these matters. This is something that can, and should, be negotiated at the "last virtual mile".

>Just because you can, doesn't mean you should and I don't feel any one entity (private or public) should be an arbiter on these matters.

What do you mean by private? Should I not be allowed to block AI agents on my sites using Cloudflare?

Re: The web does not need gatekeepers: Cloudflare’s new “signed agents” pitch

#403
post #12

I pretty much use Perplexity exclusively at this point, instead of Google. I'd rather just get my questions answered than navigate all of the ads and slowness that Google provides. I'm fine with paying a small monthly fee, but I don't want Cloudflare being the gatekeeper. Perhaps a way to serve ads through the agents would be good enough. I'd prefer that to be some open protocol than controlled by a company.

>but I don't want Cloudflare being the gatekeeper

Cloudflare is not the gatekeeper, it's the owner of the site that blocks Perplexity that's "gatekeeping" you. You're telling me that's not right?

Re: The web does not need gatekeepers: Cloudflare’s new “signed agents” pitch

#404

Earlier quoted context omitted.

> Does allow bots to access my information prevent other people from accessing my information? No. Yes it does, that's the entire point. The flood of AI bots is so bad that (mainly older) servers are literally being overloaded and (newer servers) have their hosting costs spike so high that it's unaffordable to keep the website alive. I've had to pull websites offline because badly designed & ban-evading AI scraper bo…

That's a problem with scrapers, not with AI. I'm not sure why there are way more AI scraper bots now than there were search scraper bots back when that was the new thing. However that's still an issue of scapers and rate limiting and nothing to do with wanting or not wanting AI to read your free and open content.

This whole discussion is about limiting bots and other unwanted agents, not about AI specifically (AI was just an obvious example)

Re: The web does not need gatekeepers: Cloudflare’s new “signed agents” pitch

#405

Earlier quoted context omitted.

You have a problem with badly behaved scrapers, not AI. I can't disagree with being against badly behaved scrapers. But this is neither a new problem or an interesting one from the idea of making information freely available to everyone, even rhinoceroses, assuming they are well behaved. Blocking bad actors is not the same thing as blocking AI.

The thing is that rhinoceroses aren't well-behaved. Even if some small fraction of them in theory might be well-behaved, the effort of trying to account for that is too small to bother. If 99% of rhinoceroses aren't well-behaved, the simple and correct response is to ban them all, and then maybe the nice ones can ask for a special permit. You switch from allow-by-default to block-by-default. Similarly it doesn't make…

ISPs are supposed to disconnect abusive customers. The correct thing to do is probably contact the ISP. Don't complain about scraping, complain about the DDOS (which is the actual problem and I'm increasingly beginning to believe the intent.)

Re: The web does not need gatekeepers: Cloudflare’s new “signed agents” pitch

#406

Earlier quoted context omitted.

> Everyone loves the dream of a free for all and open web... But the reality is how can someone small protect their blog or content from AI training bots? Aren't these statements entirely in conflict? You either have a free for all open web or you don't. Blocking AI training bots is not free and open for all.

No, that is not true. It is only true if you just equate "AI training bots" with "people" on some kind of nominal basis without considering how they operate in practice. It is like saying "If your grocery store is open to the public, why is it not open to this herd of rhinoceroses?" Well, the reason is because rhinoceroses are simply not going to stroll up and down the aisles and head to the checkout line quietly wit…

> "If your grocery store is open to the public, why is it not open to this herd of rhinoceroses?"

What this scenario actually reveals is that the words "open to the public" are not intended to mean "access is completely unrestricted".

It's fine to not want to give completely unrestricted access to something. What's not fine, or at least what complicates things unnecessarily, is using words like "open and free" to describe this desired actually-we-do-want-to-impose-certain-unstated-restrictions contract.

I think people use words like "open and free" to describe the actually-restricted contracts they want to have because they're often among like-minded people for whom these unstated additional restrictions are tacitly understood -- or, simply because it sounds good. But for precise communication with a diverse audience, using this kind of language is at best confusing, at worst disingenuous.

Re: The web does not need gatekeepers: Cloudflare’s new “signed agents” pitch

#407
post #16

Well, if you have a better way to solve this that’s open I’m all ears. But what Cloudflare is doing is solving the real problem of AI bots. We’ve tried to solve this problem with IP blocking and user agents, but they do not work. And this is actually how other similar problems have been solved. Certificate authorities aren’t open and yet they work just fine. Attestation providers are also not open and they work just…

> Well, if you have a better way to solve this that’s open I’m all ears. Regulation. Make it illegal to request the content of a webpage by crawler if a website operator doesn't explicitly allows it via robots.txt. Institute a government agency that is tasked with enforcement. If you as a website operator can show that traffic came from bots, you can open a complaint with the government agency and they take care of s…

The biggest issue right now seems to be people renting their residential IP addresses to scraper companies, who then distribute large scrapes across these mostly distinct IPs. These addresses are from all over the world, not just your own country, so we'll either need a World Government, or at least massive intergovernmental cooperation, for regulation to help.

Re: The web does not need gatekeepers: Cloudflare’s new “signed agents” pitch

#408

Earlier quoted context omitted.

> Well, if you have a better way to solve this that’s open I’m all ears. Regulation. Make it illegal to request the content of a webpage by crawler if a website operator doesn't explicitly allows it via robots.txt. Institute a government agency that is tasked with enforcement. If you as a website operator can show that traffic came from bots, you can open a complaint with the government agency and they take care of s…

The biggest issue right now seems to be people renting their residential IP addresses to scraper companies, who then distribute large scrapes across these mostly distinct IPs. These addresses are from all over the world, not just your own country, so we'll either need a World Government, or at least massive intergovernmental cooperation, for regulation to help.

I don't think we need a world government to make progress on that point.

The companies buying these services, are buying them from other companies. Countries or larger blocks like the EU can exert significant pressure on such companies by declaring the use of such services as illegal when interacting with websites hosted in the country or block or by companies in them.

Re: The web does not need gatekeepers: Cloudflare’s new “signed agents” pitch

#409
post #20

Everyone loves the dream of a free for all and open web. But the reality is how can someone small protect their blog or content from AI training bots? E.g.: They just blindly trust someone is sending Agent vs Training bots and super duper respecting robots.txt? Get real... Or, fine what if they do respect robots.txt, but they buy the data that may or may not have been shielded through liability layers via "licensed d…

> Everyone loves the dream of a free for all and open web... But the reality is how can someone small protect their blog or content from AI training bots? Aren't these statements entirely in conflict? You either have a free for all open web or you don't. Blocking AI training bots is not free and open for all.

Nothing is „free“. AI bots eat up my blog like crazy and I have to pay for its hosting.

Re: The web does not need gatekeepers: Cloudflare’s new “signed agents” pitch

#410
post #114

Earlier quoted context omitted.

> Wait, what? I was referring to the following image: https://substackcdn.com/image/fetch/$s_!zRK-!,w_1250,h_703,c...

I know the image, what I do not understand is the argument between using it being incompatible with "fairness" and "openness"

I can't speak for the other commenter, but I think companies like Midjourney and OpenAI are robber barons exploiting people's creative work in ways that obviously aren't fair, but that our legal system wasn't equipped to prevent.
Post reply on HN