Live data from Hacker News

Cloudflare's new AI traffic options for customers

blog.cloudflare.com

151–160 of 169 posts

Re: Cloudflare's new AI traffic options for customers

#151

Please consider installing one of the many PoW schemes such as anubis rather than use these cloudflare "features". I increasingly encounter outright blocks rather than any sort of captcha when visiting cloudflare "protected" sites. Each individual site isn't particularly important to me but it's depressing to watch the process unfold like this. You really are choosing to erode the core basis of the internet if you go…

PoW schemes like Anubis don't work. Increasingly bots are using headless browsers and are basically able to solve captchas, proof-of-work(s) and basically bypass all any any attempts to block them. It's becoming impossible to stop bots from hammering your sites/services for unwanted traffic.

Seems like there is no POW "solving" if the POW in question is literally solving some sort of mathematical puzzle. It has a bounds and we can just scale the bounds depending on how much abuse we are seeing.

I've never understood why that isn't more common honestly. It would be cool if I could actually transfer POW from one domain to another as well. Probably wouldn't stop scrapers, but it would likely stop low margin fraud type things.

Re: Cloudflare's new AI traffic options for customers

#152
post #59

Please consider installing one of the many PoW schemes such as anubis rather than use these cloudflare "features". I increasingly encounter outright blocks rather than any sort of captcha when visiting cloudflare "protected" sites. Each individual site isn't particularly important to me but it's depressing to watch the process unfold like this. You really are choosing to erode the core basis of the internet if you go…

Not sure why Anubis is getting so much hype on HN, but honestly, it is not the solution. A real solution would use behavioral modeling. Most browser fingerprinting issues are already largely solved anyway.

I'm genuinely baffled that this is your lived experience. I wonder if you only look at a very restricted set of websites?

Re: Cloudflare's new AI traffic options for customers

#153
post #86

Earlier quoted context omitted.

It's a cache. My tiny websites couldn't survive getting hammered by AI bots without them.

Are you sure? Have you tried, or did Cloudflare just tell you that?

I wouldn't need a cache if my $6 server could handle 1M hits a day.

Re: Cloudflare's new AI traffic options for customers

#154

Earlier quoted context omitted.

As things stand right now, the current web may not be possible to maintain for old and low-end phones given the costs imposed by AI training crawlers. Please propose an alternative to both Cloudflare and Anubis, that shields websites against inhuman traffic without frequent operator intervention (or otherwise negates the capacity costs they pay for AI crawling) and is compatible with low-end smartphones. Certainly, I…

Great you now made the internet unusable for half the planet.

Sounds like something the government should regulate and own then rather than relying on corporate benefactors that all profit off of malware advertising.

Re: Cloudflare's new AI traffic options for customers

#155

Earlier quoted context omitted.

Search traffic is nose diving due to LLM use. The fundamental calculus with google is you install GA and it helps your SEO has changed and google is riding out what will eventually wither as people start to reevaluate the trade off.

It's nose diving due to Google putting the LLM answer box at the top of the search results. But then isn't it even worse to be blocking every search engine crawler that isn't doing that?

Even companies like OpenAI segregate their training crawlers (GPTBot for general corpus training; OAI-SearchBot for their search index inclusion; and ChatGPT-User for user-driven agentic browser use, etc).

When even OpenAI is more respectful of intellectual property and website owner control than you, that's a problme.

Re: Cloudflare's new AI traffic options for customers

#156

Earlier quoted context omitted.

Except that officially, that is not what they do? It's a dick move not to put the two usecases under separate user agents, but their documentation says you're free to block Google-Extended via robots.txt which is used for training and grounding, while still being included in the search index. > Google-Extended does not impact a site's inclusion in Google Search nor is it used as a ranking signal in Google Search. htt…

AI crawlers, famous for respecting robots.txt ;)

The western AI crawlers are generally very-well behaved. GPTBot and ClaudeBot are nice crawlers and I see them respecting my robots.txt.

There are of course plenty of mystery, obfuscated/camouflaged scrapers/crawlers from god knows whom. Thankfully they are easy to spot and ban, although I've definitely thought about deliberating sending them poisoned data.

Re: Cloudflare's new AI traffic options for customers

#157

Earlier quoted context omitted.

Great you now made the internet unusable for half the planet.

Sounds like something the government should regulate and own then rather than relying on corporate benefactors that all profit off of malware advertising.

That will be about as effective as attempts to block soccer streams in Spain are, and I don't think it'll pan out as a viable solution without worldwide enforcement backed by BGP bans of providers that continue to host non-compliant sites after notified. It's technically possible to consider a future where RIPE might withdraw your BGP assignment if you host advertising malware on your website, but ultimately that's just as dystopic as the DMCA is today, only scaled up to worldwide impacts and drama.

Re: Cloudflare's new AI traffic options for customers

#158

Earlier quoted context omitted.

As things stand right now, the current web may not be possible to maintain for old and low-end phones given the costs imposed by AI training crawlers. Please propose an alternative to both Cloudflare and Anubis, that shields websites against inhuman traffic without frequent operator intervention (or otherwise negates the capacity costs they pay for AI crawling) and is compatible with low-end smartphones. Certainly, I…

Great you now made the internet unusable for half the planet.

AI is absolutely going to make the internet unusable for half the planet if we don't find a better way; it's happening in realtime and I hate to see it. I would love to know of a better way, but everyone's so busy clamoring against effective solutions without proposing anything else that would work.

Doing nothing doesn't work. Doing Cloudflare isn't acceptable. Doing Anubis makes the Internet unusable "for half the planet". I'm on your side morally, but site operators can't afford to be idealists in the face of AI crawling bills.

What do you, does anyone, suggest that hasn't been tried, or adopted, or considered? What's left that will defend operators against AI's traffic flood that isn't eating a thousand doller crawler bill, paying Cloudflare, or locking out half the planet?

Re: Cloudflare's new AI traffic options for customers

#159

Earlier quoted context omitted.

You are asserting that they don't work without explanation or evidence. Meanwhile it isn't clear why they wouldn't and indeed they appear to accomplish the stated goal of severely rate limiting scrapers.

IIRC there was a blog post here on HN recently explaining in detail why Anubis doesn't keep out AI bots.

Yes the default config can be trivially bypassed. When it becomes a problem the UA exceptions can be removed. It's no different than how difficulty can be scaled up and down as circumstance demands.

Re: Cloudflare's new AI traffic options for customers

#160
post #155

Earlier quoted context omitted.

It's nose diving due to Google putting the LLM answer box at the top of the search results. But then isn't it even worse to be blocking every search engine crawler that isn't doing that?

Even companies like OpenAI segregate their training crawlers (GPTBot for general corpus training; OAI-SearchBot for their search index inclusion; and ChatGPT-User for user-driven agentic browser use, etc). When even OpenAI is more respectful of intellectual property and website owner control than you, that's a problme.

> Even companies like OpenAI segregate their training crawlers (GPTBot for general corpus training; OAI-SearchBot for their search index inclusion; and ChatGPT-User for user-driven agentic browser use, etc).

They have three different user agent strings, anyway. The problem is obviously that you can't tell what someone does with data after you give it to them.

You also don't know that some third party isn't crawling the web with the user agent string "OAI-SearchBot" and then using the results for training or selling the data to the likes of Anthropic or OpenAI without telling them that.

In general attempting to use the user agent string for access control is not going to work and it seems like Google is being the less disingenuous party on this one by not blowing smoke.

Post reply on HN