Please consider installing one of the many PoW schemes such as anubis rather than use these cloudflare "features". I increasingly encounter outright blocks rather than any sort of captcha when visiting cloudflare "protected" sites. Each individual site isn't particularly important to me but it's depressing to watch the process unfold like this. You really are choosing to erode the core basis of the internet if you go…
Cloudflare's new AI traffic options for customers
41–50 of 169 posts
Re: Cloudflare's new AI traffic options for customers
#42What’s the end goal for Cloudflare and the web here? I don’t think ADOG (anthropic, deepmind, openai, google) is going to pay to crawl. What would force their hand? It’s more likely they’ll strike undisclosed agreements with major sources of discussion like reddit etc. That’s not to say getting new information as a way of context-providing is not going to happen but that’s not scraping.
Something has to happen or Google will end up starved and locked out of everything, by means both technical and legal. Then nobody gets anything.
I don't have the answer as to what happens next, and I doubt anyone else who proclaims one super confidently. But we can do some constraints analysis. There is no world where everyone works for free so Google and other AI engines can get all the value from the content, so we can eliminate those possibilities. I think we can safely discard the world(s) in which all content production just stops. However, off the top of my head, it's hard to get much tighter than that, and that definitely leaves a world where effectively everything everywhere ends up going pay-to-access.
Microtransactions have, to date, failed comprehensively, though, so the constraints on what "everything is pay-to-access" gets weird without them.
And there is never guarantee that there is any solution to any set of constraints. Things can end up overconstrained in reality as easily as a math problem. I don't actually think it'll go that way, but when analyzing this question I think it's important to not let "but $SOMETHING just has to have some way to work, because... uh... it has to!" Let the constraints do the talking. You could end up with a scenario where all content of any value is locked down, and it's fundamentally difficult and expensive to ever access or discover it, and consequently the entire content production industry radically contracts compared to its current size, if there is no pragmatic solution to microtransactions that is low-enough friction to get over the psychological and economic hurdles that have killed it to date. If everything is locked behind "macrotransactions" that's a much smaller commercial web. Probably a much higher quality one, too, but at a pretty stiff cost.
Re: Cloudflare's new AI traffic options for customers
#43The big news here is that Googlebot will be blocked from September 15th onwards by one the "block training" policies, because Google use the same crawler infrastructure for their search index AND for training Gemini: > Another change that will apply on September 15 is that multi-purpose crawlers (specifically those that combine Search with Training) will be allowed/blocked according to all of their behaviors, in line…
That's exactly what antitrust laws are supposed to do, and I hope at least EU regulators take action. Every single Googlebot crawl in your access logs is a trace for damages.
Re: Cloudflare's new AI traffic options for customers
#44What’s the end goal for Cloudflare and the web here? I don’t think ADOG (anthropic, deepmind, openai, google) is going to pay to crawl. What would force their hand? It’s more likely they’ll strike undisclosed agreements with major sources of discussion like reddit etc. That’s not to say getting new information as a way of context-providing is not going to happen but that’s not scraping.
The usage patterns of how people pay and use AI is basically the same model the web should be using: you pay a small bit of money to access monetized pages, just how you pay a small bit of money to get AI responses.
It just needs people and browsers to get onboard with protocols. Crawlers will have no choice but to pay for content behind these 402 gateways.
Re: Cloudflare's new AI traffic options for customers
#45Earlier quoted context omitted.
> And if you're getting a turnstile challenge it's unclear how it's different than an anubis challenge. You are guaranteed to pass an Anubis challenge eventually [0], whereas it's possible to get stuck forever in an infinitely-looping Turnstile challenge. > If you're outright blocked, it's probably a site decision (eg. block all VPNs or block everyone not from a given country) rather than cloudflare's. Cloudflare blo…
>You are guaranteed to pass an Anubis challenge eventually [0] That's a double edged sword because bots will eventually get through too, and unlike humans, their time is dirt cheap. >Cloudflare blocks legitimate users itself sometimes [1]. I never ran into this issue despite using seemingly maximally suspicious configs like tor browser. I can't say the same for some other vendors.
Yeah, I really have no idea why Anubis works right now: residential proxies are far more expensive than compute, yet the bots seem to have no problem obtaining millions of residential IPs, but they give up on even short-ish Anubis challenges.
> I never ran into this issue despite using seemingly maximally suspicious configs like tor browser. I can't say the same for some other vendors.
Yeah, I don't like the Cloudflare challenges, but in the past 5 years I've only had it outright block me once, and that fixed itself after 15 minutes. And I use Firefox on Linux with various privacy extensions, so my browser probably appears at least moderately suspicious.
Whereas I've been trapped in impossible ReCaptcha loops quite a few times, which is still better than vague error messages that magically go away when I switch to something not running Linux. So I'll begrudgingly accept that Turnstile is the least user-hostile product on the market right now.
Re: Cloudflare's new AI traffic options for customers
#46Please consider installing one of the many PoW schemes such as anubis rather than use these cloudflare "features". I increasingly encounter outright blocks rather than any sort of captcha when visiting cloudflare "protected" sites. Each individual site isn't particularly important to me but it's depressing to watch the process unfold like this. You really are choosing to erode the core basis of the internet if you go…
Re: Cloudflare's new AI traffic options for customers
#47> For all new domains onboarding to Cloudflare, the categories of Training and Agent will be blocked by default on the pages that display ads, while Search will remain allowed by default. It's kind of exhausting seeing Cloudflare playing both sides of the arms race. I just can't imagine bringing myself to use their technology to build agents and build AI products when they're also doing things like this. > This also…
How are they playing both sides? I thought their scraping products were also about having it behave and not take down systems
This always engenders a solid amount of distaste from me, because much like Google and Chrome, it creates the incentive for you to treat yourself better than others. Especially coupled with the trust stuff. Of course, Cloudflare is always going to trust their own platform.
Re: Cloudflare's new AI traffic options for customers
#48> For all new domains onboarding to Cloudflare, the categories of Training and Agent will be blocked by default on the pages that display ads, while Search will remain allowed by default. It's kind of exhausting seeing Cloudflare playing both sides of the arms race. I just can't imagine bringing myself to use their technology to build agents and build AI products when they're also doing things like this. > This also…
Re: Cloudflare's new AI traffic options for customers
#49What’s the end goal for Cloudflare and the web here? I don’t think ADOG (anthropic, deepmind, openai, google) is going to pay to crawl. What would force their hand? It’s more likely they’ll strike undisclosed agreements with major sources of discussion like reddit etc. That’s not to say getting new information as a way of context-providing is not going to happen but that’s not scraping.
I get a mixed feeling about all this. Cloudflare is unilaterally making all these decisions which impact the whole internet traffic flow. Taking the lead is one thing, however decisions like this should have the direct involvement of Internet Engineering Task Force (IETF) to account for all stakeholders, otherwise we run into the situation of a fragmented internet
Re: Cloudflare's new AI traffic options for customers
#50Earlier quoted context omitted.
>You are guaranteed to pass an Anubis challenge eventually [0] That's a double edged sword because bots will eventually get through too, and unlike humans, their time is dirt cheap. >Cloudflare blocks legitimate users itself sometimes [1]. I never ran into this issue despite using seemingly maximally suspicious configs like tor browser. I can't say the same for some other vendors.
> That's a double edged sword because bots will eventually get through too, and unlike humans, their time is dirt cheap. Yeah, I really have no idea why Anubis works right now: residential proxies are far more expensive than compute, yet the bots seem to have no problem obtaining millions of residential IPs, but they give up on even short-ish Anubis challenges. > I never ran into this issue despite using seemingly ma…