Earlier quoted context omitted.
You have a problem with badly behaved scrapers, not AI. I can't disagree with being against badly behaved scrapers. But this is neither a new problem or an interesting one from the idea of making information freely available to everyone, even rhinoceroses, assuming they are well behaved. Blocking bad actors is not the same thing as blocking AI.
The thing is that rhinoceroses aren't well-behaved. Even if some small fraction of them in theory might be well-behaved, the effort of trying to account for that is too small to bother. If 99% of rhinoceroses aren't well-behaved, the simple and correct response is to ban them all, and then maybe the nice ones can ask for a special permit. You switch from allow-by-default to block-by-default. Similarly it doesn't make…
The web does not need gatekeepers: Cloudflare’s new “signed agents” pitch
461–470 of 520 posts
Re: The web does not need gatekeepers: Cloudflare’s new “signed agents” pitch
#462With what they say about authorization, I think X.509 would help. (Although central certificate authorities are often used with X.509, it does not have to be that way; the service you are operating can issue the certificate to you instead, or they can accept a self-signed certificate which is associated with you the first time it is used to create an account on their service.) You can use the admin certificate issued…
What problem does this solve that a basic API key doesn't solve already? The issue with that approach is that you will require accounts/keys/certificates for all hosts you intend to visit, and malicious bots can create as many accounts as they need. You're just adding a registration step to the crawling process. Your suggested approach works for websites that want to offer AI access as a service to their customers, b…
Many things, including improved security, and the possibility of delegating authorization in the ways described in their article (if you do not restrict the certificate from issuing further certificates, and if you define an extension for use with your service to specify narrower authorization, and document this).
> The issue with that approach is that you will require accounts/keys/certificates for all hosts you intend to visit, and malicious bots can create as many accounts as they need. You're just adding a registration step to the crawling process.
Read the last paragraph of what I wrote, which explains why that issue does not apply. However, even if registration is required (which I say should not be required for most things anyways, especially read-only stuff), it does not necessarily have to be that fast or automatic.
> Your suggested approach works for websites that want to offer AI access as a service to their customers, but the problem Cloudflare is trying to solve is that most AI bots are doing things that website owners don't want them to do. The goal is to identify and block bad actors, not to make things easier for good actors.
The approach I describe would work for many things where authentication and authorization helps (most of which does not involve AI).
I do know that it does not solve the problem that Cloudflare is trying to solve, but it does what it says in the article about authorization, and in a secure way. And, it is open, interoperable, and standardized.
The problem that Cloudflare is trying to solve cannot be solved in this way, and the way Cloudflare tries to do it is not good either.
Things that AI bots are doing to other's sites includes such things as excessive scraping, rather than accessing private data (even if they might do that too, Cloudflare's solution won't help with that at all either). (There is also excessive blocking, but Cloudflare is a part of the problem, even if some of the things they do sometimes help.)
See comment 45068556. Not everything should require authentication or authorization. Also see many other comments, that also mention why it does not help.
> Using mTLS/client certificates also exposes people (that don't use AI bots) to the awful UI that browsers have for this kind of authentication. We'll need to get that sorted before an X509-based solution makes any sense.
OK, it is a valid point, but this could be improved, independently. (Before it is fixed (and even afterward if wanted), X.509 could be made as only one type of authentication; the service could allow using a username/password (and/or other things, such as TOTP) as well for people who do not want to use X.509.)
Also, AI bots are not the only kind of automated access (and is not one that I use personally, although other people might); you could also be using a API for other purposes, or you might be using a command-line program for manual access without the use of a web browser, etc.
Re: The web does not need gatekeepers: Cloudflare’s new “signed agents” pitch
#463Earlier quoted context omitted.
No, that is not true. It is only true if you just equate "AI training bots" with "people" on some kind of nominal basis without considering how they operate in practice. It is like saying "If your grocery store is open to the public, why is it not open to this herd of rhinoceroses?" Well, the reason is because rhinoceroses are simply not going to stroll up and down the aisles and head to the checkout line quietly wit…
> "If your grocery store is open to the public, why is it not open to this herd of rhinoceroses?" What this scenario actually reveals is that the words "open to the public" are not intended to mean "access is completely unrestricted". It's fine to not want to give completely unrestricted access to something. What's not fine, or at least what complicates things unnecessarily, is using words like "open and free" to des…
The other thing, though, is that there's a difference between "I personally want to release my personal work under open, free, and unrestricted terms" and "I want to release my work into a system that allows people to access information in general under open, free, and unrestricted terms". You can't just look at the individual and say "Oh, well, the conditions you want to put on your content mean it's not open and free so you must not actually want openness and freedom". You have to look at the reality of the entire system. When bots are overloading sites, when information is gated behind paywalls, when junk is firehosed out to everyone on behalf of paid advertisers while actual websites are down on page 20 of the search results, the overall situation is not one of open and free information exchange, and it's naive to think that individuals simply dumping their content "openly and freely" into this environment is going to result in an open and free situation.
Asking people to just unilaterally disarm by imposing no restrictions, while other less noble actors continue to impose all sorts of restrictions, will not produce a result that is free of restrictions. In fact quite the opposite. In order to actually get a free and open world in the large, it's not sufficient for good actors to behave in a free and open manner. Bad actors also must be actively prevented from behaving in an unfree and closed manner. Until they are, one-sided "gifts" of free and open content by the good actors will just feed the misdeeds of the bad actors.
Re: The web does not need gatekeepers: Cloudflare’s new “signed agents” pitch
#464Earlier quoted context omitted.
It should have the same protections as an EULA, where the crawler is the end user, and crawlers should be required to read it and apply it.
So none at all? EULAs are mostly just meant to intimidate you so you won't exercise your inalienable rights.
Re: The web does not need gatekeepers: Cloudflare’s new “signed agents” pitch
#465Earlier quoted context omitted.
Problem one is they do not honor the conventions of the web and abuse the sites. Problem two is they are taking content for free, distilling it into a product, and limiting access to that product.
Problem one is not specific to AI and not even about AI. Problem two is not anything new. Taking freely available content and distilling it into a product is something valuable and potentially worth paying for. People used to buy encyclopedias too. There are countless examples.
It was a similar problem with cryptocurrencies. Out comes some kind of tech thingy, and a million get-rich-quick scammers pop out of the woodwork and start scamming left, right and center. Suddenly everyone's in on the hustle, everyone's cryptomining, or taking over computers and using them for cryptomining, they're setting the world on fire with electricity consumption through the roof just to fight against other people (who they wouldn't need to fight against if they'd just cooperate).
A vision. A gold rush. A massive increase in shitty human behaviour motivated by greed.
And now here we are again with AI. Massive interest. Trillions of dollars being sloshed around, everyone hustling to develop something so they'll get picked and flooded with cash. An enormous pile of deeply unethical and disrespectful behaviour by people who are doing what they're doing because that's where the money is. The AI bubble.
Re: The web does not need gatekeepers: Cloudflare’s new “signed agents” pitch
#466I use uncommon web browsers that don't leak a lot of information. To Cloudflare, I am indistingushable from a bot. Privacy cannot exist in an environment where the host gets to decide who access the web page. I'm okay with rate limiting or otherwise blocking activity that creates too much of a load, but trying to prevent automated access is impossible withou preventing access from real people.
Honestly, just let the bots or whatever through. It's absolutely ridiculous locking out real people who did nothing wrong.
Re: The web does not need gatekeepers: Cloudflare’s new “signed agents” pitch
#467Earlier quoted context omitted.
I'm not talking about ads or pixels, I'm referring to bot operators creating so much traffic that the network bill makes the hosting financially impossible > my answer is no. Rights for me, but not for thee?
You have every right to take the content offline, or to put any technical barriers you desire in place to access it - but that's about all you should be able to do. If you don't want to lose money and don't feel confident that you can protect your content with technical measures, best to take your stuff off the internet.
Re: The web does not need gatekeepers: Cloudflare’s new “signed agents” pitch
#468Everyone loves the dream of a free for all and open web. But the reality is how can someone small protect their blog or content from AI training bots? E.g.: They just blindly trust someone is sending Agent vs Training bots and super duper respecting robots.txt? Get real... Or, fine what if they do respect robots.txt, but they buy the data that may or may not have been shielded through liability layers via "licensed d…
Re: The web does not need gatekeepers: Cloudflare’s new “signed agents” pitch
#469Earlier quoted context omitted.
If an AI bot is accessing my site the way that regular users are accessing my site -- in other words everyone is using the town center as intended -- what is the problem? Seems to be a lot of conflating of badly coded (intentionally or not) scrapers and AI. That is a problem that predates AI's existence.
So if I buy a DDoS service and DDoS your site, it's ok as long as it accesses it the same way regular people do? In sorry for extreme example, it's obviously not, but that's how I understand your position as written. We can also consider 10 exploit attempts per second that my site sees.
Re: The web does not need gatekeepers: Cloudflare’s new “signed agents” pitch
#470Earlier quoted context omitted.
But then we get to use those AI tools. The refrain here comes down not to "AI" but mostly to "the AI bot assault" which is a different thing. Sure lets have an discussion about badly behaved and overzealous web scrapers. As for credit, I've asked AI for it's references and gotten them. If my information is merely mushed into AI training model I'm not sure why I need credit. If you discuss this thread with your friend…
No, you don't "get to" use the AI tools. You have to buy access to them (beyond some free trials).