Live data from Hacker News

The web does not need gatekeepers: Cloudflare’s new “signed agents” pitch

positiveblue.substack.com

301–310 of 520 posts

Re: The web does not need gatekeepers: Cloudflare’s new “signed agents” pitch

#301

Earlier quoted context omitted.

I self-host lots of stuff. But yes it is more pain to host a WAF that can handle billions of request per minute. Even harder to do it for free like Cloudflare. And in the end the end result for the user is exactly the same if you use a self-hosted WAF or let someone else host it for you.

But you don't get billions of requests per minute. You get maybe five requests per second (300 per minute) on a bad day. The sites that seem to be getting badly attacked, they get 200 per second, which is still within reach of a self hosted firewall. Think about how many CPU cycles per packet that allows for. Hardly a real DDoS. The only reason you even want to firewall 200 requests per second is that the code downst…

Such entitlement.

How much additional tax money should I spend at work so the AI scum can make 200 searches per second?

Human and 'nice' bots make about 5 per second.

Re: The web does not need gatekeepers: Cloudflare’s new “signed agents” pitch

#303
post #278
post #276

Earlier quoted context omitted.

The dream is real, man. If you want open content on the Internet, it's never been a better time. My blog is open to all - machine or man. And it's hosted on my home server next to me. I don't see why anyone would bother trying to distinguish humans from AI. A human hitting your website too much is no different from an AI hitting your website too much. I have a robots.txt that tries to help bots not get stuck in loops…

It's traditional to include a link when claiming to be invulnerable. :)

Haha, sounds a bit self-promotional to do that but link in profile.

Not claiming that the site is technologically invulnerable. Just that it's not a big deal if LLMs scrape it (which bizarrely they do).

Re: The web does not need gatekeepers: Cloudflare’s new “signed agents” pitch

#304

Cloudflare slows the whole damn websites down. It takes many seconds to deal with their trash. I hope they crash and burn. Let's get back to very low latency websites without the cloudflare garbage.

Cloudflare as a CDN greatly greatly speeds up the web.

All the custom code they write on top of that to transform HTML for you? Ehhhh... don't use those features. Most are easily reproducible on the backend.

Re: The web does not need gatekeepers: Cloudflare’s new “signed agents” pitch

#305
post #270

Earlier quoted context omitted.

> Both whitelist and allowlist are equally normal and good. Good is debatable, but normal ?? No, obviously not. One is a word that has been around for over a hundred years and is understood by everyone that speaks English; the other is like 4 years old and only used by some software nerds. 90% of normal people would not know what it means.

90% of normal people have never heard of a whitelist either, but any English speaker could intuit what "allowlist" means more easily that "whitelist" without context. And both are technical terms of art in the context of this conversation, so what 90% of normal people would or wouldn't understand isn't even relevant. And "enshittification" is even newer than "allowlist" and it's practically mainstream.

> any English speaker could intuit what "allowlist" means more easily that "whitelist" without context

That is not true IMO. Blacklist is a standard English word that any native speaker would know; whitelist (while not as standard) is easy to extrapolate from that.

Re: The web does not need gatekeepers: Cloudflare’s new “signed agents” pitch

#306
post #20

Everyone loves the dream of a free for all and open web. But the reality is how can someone small protect their blog or content from AI training bots? E.g.: They just blindly trust someone is sending Agent vs Training bots and super duper respecting robots.txt? Get real... Or, fine what if they do respect robots.txt, but they buy the data that may or may not have been shielded through liability layers via "licensed d…

> Everyone loves the dream of a free for all and open web. But the reality is how can someone small protect their blog or content from AI training bots? I'm old enough to remember when people asked the same questions of Hotbot, Lycos, Altavista, Ask Jeeves, and -- eventually -- Google. Then, as now, it never felt like the right way to frame the question. If you want your content freely available, make it freely avail…

Google (and the others) crawl from a published IP range, with "Google" in the user agent. They read robots.txt. They are very easy to block

The AI scum companies crawl from infected botnet IPs, with the user agent the same as the latest Chrome or Safari.

Re: The web does not need gatekeepers: Cloudflare’s new “signed agents” pitch

#307

Earlier quoted context omitted.

Allowlist is arguably fitting for a list of things which are allowed.

There are so many terms in software which are nonsensical (starting with "computer science") which could be fixed. The problem with changing whitelist to "allowlist" is that it implies that people who use whitelist are racists. You're not just virtue signaling (and confusing my spellchecker) but causing discord. It would be perfectly fine if people switched to "allowlist" because they think it's a better term, but th…

That is exactly why I hate "allowlist", "main" instead of "master", and so on. The reason they were proposed is because some people were trying to play dominance games with grievance politics. We should attempt to resist such bad faith tactics, not propagate them. And yes, unfortunately that means I have to take a stand on something that is otherwise inconsequential. But such is the price of pushing back on self-righteous prigs who are trying to police terms of art.

Re: The web does not need gatekeepers: Cloudflare’s new “signed agents” pitch

#308

Earlier quoted context omitted.

Creative Commons, GFDL, Unlicense, GPL/AGPL, MIT, WTFPL. Go crazy. I have the freedom to police how users use the information on my site. Yes. Real examples: My blog is BY-NC-SA and digital garden is GFDL. You can't take them, mangle and sell them. Especially, the blog. AI companies take these posts, and sell derivatives, without any references, consent or compensation. BY-NC-SA is complete opposite of what they do.…

Absolutely. If you want to put all kinds of copyright, license, and even payment restrictions on your content go ahead. And if AI companies or people abuse that, that's bad on them. But I do think if you're serious about free and open information than why are you doing that in the first place? It's perfectly reasonable to be restrictive; I write both very open software and very closed software. But I see a lot of peo…

Let me try to make my point as compact as possible. I may fail, but please bear with me.

I prefer Free Software to Open Source software. My license of choice is A/GPLv3+. Because, I don't want my work to be used by people/entities in a single sided way. The software I put out is the software I develop for myself, with the hope of being useful for somebody else. My digital garden is the same. My blog is a personal diary in the open. These are built on my free time, for myself, and shared.

See, permissive licenses are for "developer freedom". You can do whatever you do with what you can grab, as long as you write a line to credits. A/GPL family is different. Wants reciprocity. It empowers the user vs. the developer. You have to give the source. Who modifies the source, shares the modifications. It stays in the open. It has to stay open.

I demand this reciprocity for what I put out there. The licenses reflect that. It's "restricting the use to keep the information/code open". I share something I spent my time on, and I want it to live on the open, want a little respect for putting out what I did. That respect is not fame or superiority. Just not take it and run with it, keeping all the improvements to yourself.

It's not yours, but ours. You can't keep it to yourself.

When it comes to AI, it's an extension of this thinking. I do not give consent to a faceless corporation to close, twist and earn money from what I put out for public good. I don't want a set of corporations act as a middleman to get what I put out, repackage and corrupt it in the process and sell it. It's not about money; it's about ethics, doing the right thing and being respectful. It's about exploitation. Same is applicable to my photos.

I'm not against AI/LLM/Generative technology/etc. I'm against exploitation of people, artists, musicians, software developers, other companies. I equally get angry when a company's source available code is scraped and used for suggestions as well as an academic's LGPL high performance matrix library which is developed via grants over the years. This thing affect livelihoods of people.

I get angry when people say "if we take permission for what we do, AI industry will collapse", or "this thing just learns like humans, this is fair use".

I don't buy their "we're doing something awesome, we need no permission" attitude. No, you need permission to use my content. Because I say so. Read the fine print.

I don't want knowledge to be monopolized by these corporations. I don't want the small fish to be eaten by the bigger one and what remains is buried into the depths of information ocean.

This is why I stopped sharing my photos for now, and my latest research won't be open source for quite some time.

What I put out is for humans' direct consumption. Middlemen are not welcome.

If you have any questions or left any holes up there, please let me know.

Re: The web does not need gatekeepers: Cloudflare’s new “signed agents” pitch

#310
post #238

Earlier quoted context omitted.

Basic damn rate limiting is pretty damn exploitable. Even ignoring botnets (which is impossible), usefully rate limiting IPv6 is anything but basic. If you just pick some prefix from /48 to /64 to key your rate limits on, you'll either be exploitable by IPs from providers that hand out /48s like candy or you'll bucket a ton of mobile users together for a single rate limit.

You make unauthenticated requests cheap enough that you don't care about volume. Reserve rate limiting for authenticated users where you have real identity. The open web survives by being genuinely free to serve, not by trying to guess who's "real." A basic Varnish setup should get you most of the way there, no agent signing required!

I guess you should start a Cloudflare competitor that just puts a cheap Varnish VM in front of websites to solve bots forever.
Post reply on HN