Live data from Hacker News

Cloudflare's new AI traffic options for customers

blog.cloudflare.com

131–140 of 169 posts

Re: Cloudflare's new AI traffic options for customers

#131

Earlier quoted context omitted.

PoW is a reasonable solution as a fallback when other metrics flag a client.

How? First, they can solve the Anubis challenge with native code, so they can solve them faster than genuine users. Second, the cost is nothing compared to training LLMs, plus they will just move the work to the residential proxies that they have access to (so, someone is paying through their TV's electricity bill). Anubis only work(s|ed) great for a while when crawlers were not prepared for these challenges. Securit…

> they will just move the work to the residential proxies that they have access to

The proxies I am familiar with do not offer arbitrary code execution. I think you're thinking of a botnet.

Regarding native code, the current crop of solutions seem to work well enough for now. Ultimately a challenge response protocol should be standardized and browsers should ship a native implementation. In the meantime WASM likely gets you close enough to native.

Re: Cloudflare's new AI traffic options for customers

#132

Earlier quoted context omitted.

How? First, they can solve the Anubis challenge with native code, so they can solve them faster than genuine users. Second, the cost is nothing compared to training LLMs, plus they will just move the work to the residential proxies that they have access to (so, someone is paying through their TV's electricity bill). Anubis only work(s|ed) great for a while when crawlers were not prepared for these challenges. Securit…

> they will just move the work to the residential proxies that they have access to The proxies I am familiar with do not offer arbitrary code execution. I think you're thinking of a botnet. Regarding native code, the current crop of solutions seem to work well enough for now. Ultimately a challenge response protocol should be standardized and browsers should ship a native implementation. In the meantime WASM likely g…

Regarding native code, the current crop of solutions seem to work well enough for now.

I just quoted a toot in another submission, adding it here since it is relevant:

We apologize for a period of extreme slowness today. The army of AI crawlers just leveled up and hit us very badly. [...] It seems like the AI crawlers learned how to solve the Anubis challenges. [...] However, we can confirm that at least Huawei networks now send the challenge responses and they actually do seem to take a few seconds to actually compute the answers. It looks plausible, so we assume that AI crawlers leveled up their computing power to emulate more of real browser behaviour to bypass the diversity of challenges that platform enabled to avoid the bot army.

https://social.anoxinon.de/@Codeberg/115033790447125787

Re: Cloudflare's new AI traffic options for customers

#133

Please consider installing one of the many PoW schemes such as anubis rather than use these cloudflare "features". I increasingly encounter outright blocks rather than any sort of captcha when visiting cloudflare "protected" sites. Each individual site isn't particularly important to me but it's depressing to watch the process unfold like this. You really are choosing to erode the core basis of the internet if you go…

Please don't use Anubis, it makes visiting websites very difficult (often multiple minutes wait times) on old hardware and low-end smartphones.

what kind of hardware would take minutes to complete the checks?

Re: Cloudflare's new AI traffic options for customers

#134
post #122

Earlier quoted context omitted.

If a site cannot handle traffic from the real Googlebot that is a serious issue with the site itself since it's actually pretty conservative Also I should note there are lots of fake Googlebots...

It is mostly, but it doesn't take a lot of googling to find sites getting ridiculous amounts of traffic from googlebot on google IPs. Its one of the most common complaints about google's search indexing

Lots of people abuse Google Cloud to get a "Google IP" for a fake Googlebot. Why don't you show me a single screenshot from Google Search Console showing a high number of requests to a site that would be counted as a DoS? All requests from the official Googlebot are logged there so if it's such a common problem it must be very easy for you to show me this.

Re: Cloudflare's new AI traffic options for customers

#135
post #43
post #4

The big news here is that Googlebot will be blocked from September 15th onwards by one the "block training" policies, because Google use the same crawler infrastructure for their search index AND for training Gemini: > Another change that will apply on September 15 is that multi-purpose crawlers (specifically those that combine Search with Training) will be allowed/blocked according to all of their behaviors, in line…

Good. Google's approach here is manifestly predator, unfair, and IMO illegal. They deserve to be in court for this behaviour, and mandating owners give consent for AI training or drop out of Google; which is just a non-starter because they're a search monopoly. That's exactly what antitrust laws are supposed to do, and I hope at least EU regulators take action. Every single Googlebot crawl in your access logs is a tr…

Except that officially, that is not what they do? It's a dick move not to put the two usecases under separate user agents, but their documentation says you're free to block Google-Extended via robots.txt which is used for training and grounding, while still being included in the search index.

> Google-Extended does not impact a site's inclusion in Google Search nor is it used as a ranking signal in Google Search. https://developers.google.com/crawling/docs/crawlers-fetcher...

Exclusion from grounding does mean that your site won't get sourced in the AI overview, but I'm not sure what the click through rates are like on those.

Re: Cloudflare's new AI traffic options for customers

#136
post #43

Earlier quoted context omitted.

Good. Google's approach here is manifestly predator, unfair, and IMO illegal. They deserve to be in court for this behaviour, and mandating owners give consent for AI training or drop out of Google; which is just a non-starter because they're a search monopoly. That's exactly what antitrust laws are supposed to do, and I hope at least EU regulators take action. Every single Googlebot crawl in your access logs is a tr…

Except that officially, that is not what they do? It's a dick move not to put the two usecases under separate user agents, but their documentation says you're free to block Google-Extended via robots.txt which is used for training and grounding, while still being included in the search index. > Google-Extended does not impact a site's inclusion in Google Search nor is it used as a ranking signal in Google Search. htt…

AI crawlers, famous for respecting robots.txt ;)

Re: Cloudflare's new AI traffic options for customers

#137
post #5

Earlier quoted context omitted.

We had googlebot blast a random customer system and almost cause an outage, this is when I first learnt that google will use it for AI training also. It's honestly kind of frustrating also because you then search on it and theres (was) nothing on how you are meant to "correctly" tell google to fuck off, and not use it like that.

The vast majority of Googlebot user agents are lying. Real Googlebot is pretty well behaved in my experience. You should use reverse DNS or IP lists to check: https://developers.google.com/crawling/docs/crawlers-fetcher...

If you're using Cloudflare, set up a security rule to block requests that have "Googlebot" in the UA and are not recognised by CF as a real bot.

Re: Cloudflare's new AI traffic options for customers

#138

What’s the end goal for Cloudflare and the web here? I don’t think ADOG (anthropic, deepmind, openai, google) is going to pay to crawl. What would force their hand? It’s more likely they’ll strike undisclosed agreements with major sources of discussion like reddit etc. That’s not to say getting new information as a way of context-providing is not going to happen but that’s not scraping.

*AGOD

Re: Cloudflare's new AI traffic options for customers

#139
post #98
post #42

Earlier quoted context omitted.

I think the best answer is, nobody knows. The previous equilibrium for content scraping for search engines on the internet was already at times an uncomfortable one. But I agree that from a game theory perspective, "the AI bots take and give nothing back in return" is not just hyperbole, it's the actual situation. If Google is successful in what seems to be its plans and it becomes a box where you type a question and…

The internet becomes full of free propaganda since it's not the consumer who pays for that?

That's the current state of the internet. In this world it wouldn't be "the Internet" full of propaganda, it would be the AI search engines the propaganda would get concentrated into. For "national security", don't you know. And it would be much easier for them for having an even smaller target. At least the current internet lets you cross-check one source of free propaganda against another today. (Although whether the truth is in any meaningful way "between" any set of them is another question.)

One possible scenario is that that is simply it for the internet as an information source; the search engine's AIs get captured and there becomes effectively no way to discover any of the content already on there.

But then again, people will react to that and do something. Kagi would grow and others too. The more interesting question is whether the governments that captured Google's AI would let them or if suddenly it would be discovered that copyright law doesn't permit search engines to do that. Would that be inconsistent with letting the Approved AIs access whatever they want and chew on it even harder? Yeah, and they wouldn't care.

Re: Cloudflare's new AI traffic options for customers

#140
post #25

Please consider installing one of the many PoW schemes such as anubis rather than use these cloudflare "features". I increasingly encounter outright blocks rather than any sort of captcha when visiting cloudflare "protected" sites. Each individual site isn't particularly important to me but it's depressing to watch the process unfold like this. You really are choosing to erode the core basis of the internet if you go…

>I increasingly encounter outright blocks rather than any sort of captcha when visiting cloudflare "protected" sites. ??? Unless you're browsing around with the googlebot user agent string, you should be getting turnstile challanges at most, not blocks. And if you're getting a turnstile challenge it's unclear how it's different than an anubis challenge. If you're outright blocked, it's probably a site decision (eg. b…

I’ve been blocked by Turnstile without possibility of challenge, no way to access nor know why. And it was with standard iOS and Windows machines and no VPN or anything. It happens.
Post reply on HN