Live data from Hacker News

The web does not need gatekeepers: Cloudflare’s new “signed agents” pitch

positiveblue.substack.com

341–350 of 520 posts

Re: The web does not need gatekeepers: Cloudflare’s new “signed agents” pitch

#342

Earlier quoted context omitted.

At present, problem one is almost entirely AI companies.

And a few decades ago, it would have been search engine scrapers instead.

And that problem was largely solved by robots.txt. AI scrapers are ignoring robots.txt and beating the hell out of sites. Small sites that have decades worth of quality information are suffering the most. Many of the scrapers are taking extreme measures to avoid being blocked, like using large numbers of distinct IP addresses (perhaps using botnets).

Re: The web does not need gatekeepers: Cloudflare’s new “signed agents” pitch

#343

Earlier quoted context omitted.

Let me try to make my point as compact as possible. I may fail, but please bear with me. I prefer Free Software to Open Source software. My license of choice is A/GPLv3+. Because, I don't want my work to be used by people/entities in a single sided way. The software I put out is the software I develop for myself, with the hope of being useful for somebody else. My digital garden is the same. My blog is a personal dia…

I respect the desire for reciprocity, but strong copyleft isn't the only, or even the best, way to protect user freedom or public knowledge. My opinion is that permissive licensing and open access to learn from public materials have created enormous value precisely because they don't pre-empt future uses. Requiring permission for every new kind of reuse (including ML training) shrinks the commons, entrenches incumben…

AGPL doesn't pre-empt future uses or require permission for any kind of re-use. You just have to share alike. It's pretty simple.

AGPL lets you take a bunch of data and AI-train on it. You just have to release the data and source code to anyone who uses the model. Pretty simple. You don't have to rent them a bunch of GPUs.

Actually it can be annoying because of the specific mechanism by which you have to share alike - the program has to have a link to its own source code - you can't just offer the source alongside the binary. But it's doable.

Re: The web does not need gatekeepers: Cloudflare’s new “signed agents” pitch

#344
post #77

Earlier quoted context omitted.

AI poisoning is a better protection. Cloudflare is capable of serving stashes of bad data to AI bots as protective barrier to their clients.

AI poisoning is going to get a lot of people killed, be cause the AI won't stop being used.

Okay, let them

Re: The web does not need gatekeepers: Cloudflare’s new “signed agents” pitch

#345

Earlier quoted context omitted.

> If you don't want AI bots reading information on the web, you don't actually want a free and open web. This is such a bad faith argument. We want a town center for the whole community to enjoy! What, you don't like those people shooting up drugs over there? But they're enjoying it too, this is what you wanted right? They're not harming you by doing their drugs. Everyone is enjoying it!

Set aside that there's a pretty big difference between AI scraping and illegal drug usage. If the person using illegal drugs is on no way harming anyone but themselves and not being a nuisance, then yeah, I can get behind that. Put whatever you want in your body, just don't let it negatively impact anyone around you. Seems reasonable?

I think this is actually a good example despite how stark the differences are - both the nuisance AI scrapers and the drug addicts have negative externalities that while possible for them to self regulate, they are for whatever reasons proving unable to do so, and therefore cause other people to have a bad time.

Other commenters saying the usual “drugs are freedom” type opinions, but now having lived in China and Japan where drugs are dealt with very strictly (and basically don’t have a drug problem today), I can see the other side of the argument where in fact places feeling dirty and dangerous because of drugs - even if you think of addicts sympathetically as victims who need help - makes everyone else less free to live the lifestyle they would like to have.

More freedom for one group (whether to ruin their own lives for a high; or to train their AI models) can mean less freedom for others (whether to not feel safe walking in public streets; or to publish their little blog in the public internet).

Re: The web does not need gatekeepers: Cloudflare’s new “signed agents” pitch

#346
post #276

Earlier quoted context omitted.

The dream is real, man. If you want open content on the Internet, it's never been a better time. My blog is open to all - machine or man. And it's hosted on my home server next to me. I don't see why anyone would bother trying to distinguish humans from AI. A human hitting your website too much is no different from an AI hitting your website too much. I have a robots.txt that tries to help bots not get stuck in loops…

> I don't see why anyone would bother trying to distinguish humans from AI. Because a hundred thousand people reading a blog post is more beneficial to the world than an AI scraper bot fetching my (unchanged) blog post a hundred thousand times just in case it's changed in the last hour. If AI bots were well-behaved, maintained a consistent user agent, used consistent IP subnets, and respected robots.txt, I wouldn't h…

I've not seen an AI scraper reading a blog post 100,000 times in an hour to see if it's changed. As far as I can tell, that's a NI hallucination. Typical fetch rates are more like 3 times per second (10k per hour) and fetch a different URL each time.

Re: The web does not need gatekeepers: Cloudflare’s new “signed agents” pitch

#347

Earlier quoted context omitted.

> Everyone loves the dream of a free for all and open web. But the reality is how can someone small protect their blog or content from AI training bots? I'm old enough to remember when people asked the same questions of Hotbot, Lycos, Altavista, Ask Jeeves, and -- eventually -- Google. Then, as now, it never felt like the right way to frame the question. If you want your content freely available, make it freely avail…

What if I want my content freely available to humans, and not to bots? Why is that such an insane, unworkable ask? All I want is a copyleft protection that specifically allows humans to access my work to their heart's content, but disallows AI use of it in any form. Is that truly so unreasonable?

Yes, it is an unreasonable and absurd ask. You cannot want freedom while restricting it. You forget that it is people that use AI agents, essentially, being cyborgs. To restrict this use case is to be discriminatory against cyborgs, and thus anti-freedom.

We are lucky that there is no way to detect it.

Re: The web does not need gatekeepers: Cloudflare’s new “signed agents” pitch

#348
post #20

Everyone loves the dream of a free for all and open web. But the reality is how can someone small protect their blog or content from AI training bots? E.g.: They just blindly trust someone is sending Agent vs Training bots and super duper respecting robots.txt? Get real... Or, fine what if they do respect robots.txt, but they buy the data that may or may not have been shielded through liability layers via "licensed d…

[deleted]

Re: The web does not need gatekeepers: Cloudflare’s new “signed agents” pitch

#349

Earlier quoted context omitted.

> Everyone loves the dream of a free for all and open web... But the reality is how can someone small protect their blog or content from AI training bots? Aren't these statements entirely in conflict? You either have a free for all open web or you don't. Blocking AI training bots is not free and open for all.

That's a very "BSD is freedom and GPL isn't" kind of philosophy. Nothing is truly free unless you give equal respect to fellow hobbyists and megacorps using your labor for their profit.

GPL doesn't care if you use it for profit or not (good), it just says that the resultant model needs to be open too. And open models exist in droves nowadays. Even closed models can be distilled into open ones.
Post reply on HN