Live data from Hacker News

Cloudflare's new AI traffic options for customers

blog.cloudflare.com

71–80 of 169 posts

Re: Cloudflare's new AI traffic options for customers

#71
post #52

>So, instead of defining a bot primarily as “AI” or not, our updated approach to classification will ask deeper questions about bot or agent behavior: What are they doing on my site? What are they storing? And how will they reshare my content? I dont get this. The question is are they a bot or a human. It doesnt matter what they are doing I dont want bots on my site.

> I dont get this. The question is are they a bot or a human. It doesnt matter what they are doing I dont want bots on my site.

Do you want your site to be discoverable by a search engine? (How do you think that occurs?)

Re: Cloudflare's new AI traffic options for customers

#72
So is it possible to say "No bots except Google, OpenAI, Grok, Claude and Perplexity"?

As far as I can tell, Google is the only one sending me visitors. And the other big AI players might do so in the future.

Another option would be "No anonymous bots". So at least if a bot would want to crawl my site, they would have to identify themselves. Since the rise of the AI bots, I am getting hurt badly with insane amounts of requests from residential IPs that mimic real humans. The only difference being they don't make me any money. Only produce costs.

By the way, how is the situation over at Amazon's Cloudfront? Do they offer something that helps? Anyone here with them?

Re: Cloudflare's new AI traffic options for customers

#73
post #43
post #4

The big news here is that Googlebot will be blocked from September 15th onwards by one the "block training" policies, because Google use the same crawler infrastructure for their search index AND for training Gemini: > Another change that will apply on September 15 is that multi-purpose crawlers (specifically those that combine Search with Training) will be allowed/blocked according to all of their behaviors, in line…

Good. Google's approach here is manifestly predator, unfair, and IMO illegal. They deserve to be in court for this behaviour, and mandating owners give consent for AI training or drop out of Google; which is just a non-starter because they're a search monopoly. That's exactly what antitrust laws are supposed to do, and I hope at least EU regulators take action. Every single Googlebot crawl in your access logs is a tr…

I don't think it's the Google bot DDOSing people's infrastructure for AI training...

Re: Cloudflare's new AI traffic options for customers

#74
post #47

Earlier quoted context omitted.

I mean that they're telling developers that they should use Cloudflare's platform to build agents, the kind of agents that would go across the web and act on behalf of users... but then they're also the ones blocking those requests. This always engenders a solid amount of distaste from me, because much like Google and Chrome, it creates the incentive for you to treat yourself better than others. Especially coupled wi…

I thought the issue people have with AI scrapers is the ones that DDoS sites to scrape everything for training purposes, rather than the ones that interactively query specific content to support an active conversation?

First one is problem because it brings no traffic back and occasionally DDoSes the site

Second one is problem because it DDoSes the site. "Just" queries for hundred thousand people (if you happened to be good source for that bit of knowledge) that don't bring actual people to your site is also a problem

Re: Cloudflare's new AI traffic options for customers

#75
post #37

Earlier quoted context omitted.

Why should I use something other than Cloudflare pages for a simple app landing page?

Because your viewer/customer base will be reduced.

It would have a domain. It affects it even then?

Re: Cloudflare's new AI traffic options for customers

#76
post #47

Earlier quoted context omitted.

I mean that they're telling developers that they should use Cloudflare's platform to build agents, the kind of agents that would go across the web and act on behalf of users... but then they're also the ones blocking those requests. This always engenders a solid amount of distaste from me, because much like Google and Chrome, it creates the incentive for you to treat yourself better than others. Especially coupled wi…

I thought the issue people have with AI scrapers is the ones that DDoS sites to scrape everything for training purposes, rather than the ones that interactively query specific content to support an active conversation?

If you read the article, you'll notice that they're explicitly making sure to block the ones that interactively query specific content too.

Re: Cloudflare's new AI traffic options for customers

#77

What’s the end goal for Cloudflare and the web here? I don’t think ADOG (anthropic, deepmind, openai, google) is going to pay to crawl. What would force their hand? It’s more likely they’ll strike undisclosed agreements with major sources of discussion like reddit etc. That’s not to say getting new information as a way of context-providing is not going to happen but that’s not scraping.

Universal tax collector of the internet. A penny for every page access. ADOG will not mind as it cements their incumbent status and pulls up the drawbridge by erecting a huge financial barrier for any new entrant.

Re: Cloudflare's new AI traffic options for customers

#78
post #18

This is fine so long as it’s easy for me to turn off. I just don’t want to accidentally lose all AI traffic one day.

Most of the internet unfortunatly ploinks fown a 'free' service and never even looks at what it defaults to. E.g., quite a lott of cloudflare "protected" sites block their rss feeds from being read by machine.

Re: Cloudflare's new AI traffic options for customers

#79
post #43

Earlier quoted context omitted.

Good. Google's approach here is manifestly predator, unfair, and IMO illegal. They deserve to be in court for this behaviour, and mandating owners give consent for AI training or drop out of Google; which is just a non-starter because they're a search monopoly. That's exactly what antitrust laws are supposed to do, and I hope at least EU regulators take action. Every single Googlebot crawl in your access logs is a tr…

I don't think it's the Google bot DDOSing people's infrastructure for AI training...

[deleted]

Re: Cloudflare's new AI traffic options for customers

#80
post #52

>So, instead of defining a bot primarily as “AI” or not, our updated approach to classification will ask deeper questions about bot or agent behavior: What are they doing on my site? What are they storing? And how will they reshare my content? I dont get this. The question is are they a bot or a human. It doesnt matter what they are doing I dont want bots on my site.

> I dont get this. The question is are they a bot or a human. It doesnt matter what they are doing I dont want bots on my site. Do you want your site to be discoverable by a search engine? (How do you think that occurs?)

Let's also spare a moment for accessibility issues. For example, is it that wrong for a blind person to invoke a tool that describes a picture when there's no alt-text? Or something which describes/transcribes audio for the deaf?

Where do we draw the line between a personal-bot and a custom browser?

Post reply on HN