Live data from Hacker News

Cloudflare's new AI traffic options for customers

blog.cloudflare.com

121–130 of 169 posts

Re: Cloudflare's new AI traffic options for customers

#121
post #43
post #4

The big news here is that Googlebot will be blocked from September 15th onwards by one the "block training" policies, because Google use the same crawler infrastructure for their search index AND for training Gemini: > Another change that will apply on September 15 is that multi-purpose crawlers (specifically those that combine Search with Training) will be allowed/blocked according to all of their behaviors, in line…

Good. Google's approach here is manifestly predator, unfair, and IMO illegal. They deserve to be in court for this behaviour, and mandating owners give consent for AI training or drop out of Google; which is just a non-starter because they're a search monopoly. That's exactly what antitrust laws are supposed to do, and I hope at least EU regulators take action. Every single Googlebot crawl in your access logs is a tr…

EU will write some strongly worded letter.

Saying this as a European who is pro EU.

Why should they?

Re: Cloudflare's new AI traffic options for customers

#122
post #15

Earlier quoted context omitted.

Google's web scraping functionality has been acting as a ddos for more than two decades. I've seen literally hundreds of reports of them attacking websites and taking them down, where there's nothing you can do but accept the traffic, or get delisted This is unfortunately nothing new. There's no correct way to tell them to fuck off, they do not care, and they never will do. People have even taken them to court over t…

If a site cannot handle traffic from the real Googlebot that is a serious issue with the site itself since it's actually pretty conservative Also I should note there are lots of fake Googlebots...

It is mostly, but it doesn't take a lot of googling to find sites getting ridiculous amounts of traffic from googlebot on google IPs. Its one of the most common complaints about google's search indexing

Re: Cloudflare's new AI traffic options for customers

#123

Please consider installing one of the many PoW schemes such as anubis rather than use these cloudflare "features". I increasingly encounter outright blocks rather than any sort of captcha when visiting cloudflare "protected" sites. Each individual site isn't particularly important to me but it's depressing to watch the process unfold like this. You really are choosing to erode the core basis of the internet if you go…

Please don't use Anubis, it makes visiting websites very difficult (often multiple minutes wait times) on old hardware and low-end smartphones.

As things stand right now, the current web may not be possible to maintain for old and low-end phones given the costs imposed by AI training crawlers.

Please propose an alternative to both Cloudflare and Anubis, that shields websites against inhuman traffic without frequent operator intervention (or otherwise negates the capacity costs they pay for AI crawling) and is compatible with low-end smartphones.

Certainly, I imagine Anubis would be interested in adopting it if it’s effective!

Re: Cloudflare's new AI traffic options for customers

#125
post #43
post #4

The big news here is that Googlebot will be blocked from September 15th onwards by one the "block training" policies, because Google use the same crawler infrastructure for their search index AND for training Gemini: > Another change that will apply on September 15 is that multi-purpose crawlers (specifically those that combine Search with Training) will be allowed/blocked according to all of their behaviors, in line…

Good. Google's approach here is manifestly predator, unfair, and IMO illegal. They deserve to be in court for this behaviour, and mandating owners give consent for AI training or drop out of Google; which is just a non-starter because they're a search monopoly. That's exactly what antitrust laws are supposed to do, and I hope at least EU regulators take action. Every single Googlebot crawl in your access logs is a tr…

This had me in disbelief since the minute I saw it: Google's "AI overview" presumably trained on content from other websites, disincentivizes users from clicking through to those websites..

How is that not conflict of interest??

Re: Cloudflare's new AI traffic options for customers

#126

Please consider installing one of the many PoW schemes such as anubis rather than use these cloudflare "features". I increasingly encounter outright blocks rather than any sort of captcha when visiting cloudflare "protected" sites. Each individual site isn't particularly important to me but it's depressing to watch the process unfold like this. You really are choosing to erode the core basis of the internet if you go…

Please don't use Anubis, it makes visiting websites very difficult (often multiple minutes wait times) on old hardware and low-end smartphones.

Personally I think that's a better outcome than the alternative.

However another option is to support both. Have a challenge page that requires the visitor to select one of several options. Cloudflare could be one of those.

Re: Cloudflare's new AI traffic options for customers

#127
post #70

Earlier quoted context omitted.

You are asserting that they don't work without explanation or evidence. Meanwhile it isn't clear why they wouldn't and indeed they appear to accomplish the stated goal of severely rate limiting scrapers.

> You are asserting that they don't work without explanation or evidence. Oh, no, that isn't how it works — you made a claim first (albeit indirectly), with no evidence. It's on you

That's not how it works. There are cases where established norms or common sense dictate a certain default assumption. This is one of them.

Still, I'll humor your absurd request by reminding you of the many success stories that have repeatedly made the front page of HN.

Beyond that we have cryptocurrencies. If you have knowledge of a generalized solution for defeating PoW schemes then why are you posting here instead of making yourself a billionaire?

Re: Cloudflare's new AI traffic options for customers

#128

Earlier quoted context omitted.

PoW schemes like Anubis don't work. Increasingly bots are using headless browsers and are basically able to solve captchas, proof-of-work(s) and basically bypass all any any attempts to block them. It's becoming impossible to stop bots from hammering your sites/services for unwanted traffic.

You are asserting that they don't work without explanation or evidence. Meanwhile it isn't clear why they wouldn't and indeed they appear to accomplish the stated goal of severely rate limiting scrapers.

IIRC there was a blog post here on HN recently explaining in detail why Anubis doesn't keep out AI bots.

Re: Cloudflare's new AI traffic options for customers

#129

Please consider installing one of the many PoW schemes such as anubis rather than use these cloudflare "features". I increasingly encounter outright blocks rather than any sort of captcha when visiting cloudflare "protected" sites. Each individual site isn't particularly important to me but it's depressing to watch the process unfold like this. You really are choosing to erode the core basis of the internet if you go…

Please don't use Anubis, it makes visiting websites very difficult (often multiple minutes wait times) on old hardware and low-end smartphones.

What's the alternative? We're stuck, either cloudflare or anubis

Re: Cloudflare's new AI traffic options for customers

#130

Earlier quoted context omitted.

You are asserting that they don't work without explanation or evidence. Meanwhile it isn't clear why they wouldn't and indeed they appear to accomplish the stated goal of severely rate limiting scrapers.

Surely I don't need to explain how headless browsers work, bot proxies and distributed crawlers to you do I? If you haven't experienced this first-hand, I can understand. Maybe it warrants a blog post. But I can assure you, these don't work at scale, they might only stop the "less sophisticated" bots and maybe (just maybe) some unwanted spam.

What do the browsers being headless or the traffic proxied have to do with anything? Each client still has to solve a PoW challenge regardless. That mitigates the (possibly unintentional) DoS by making it exorbitantly expensive.
Post reply on HN