Live data from Hacker News

Cloudflare CEO is lying to you about the bot traffic jump

flyingpenguin.com

81–90 of 149 posts

Re: Cloudflare CEO is lying to you about the bot traffic jump

#81

I deal with scrapers that sometimes border on DDoSes for LessWrong. The amount of bot traffic varies greatly between sites; if you have more URLs you get more bot traffic (regardless of whether those URLs represent a deep content catalog, or useless URL parameter permutations). It's bad for LW because of the content-catalog depth. It's easy to drastically underestimate the amount of bot traffic, because bots make eff…

[deleted]

Re: Cloudflare CEO is lying to you about the bot traffic jump

#82
post #48
post #46

Earlier quoted context omitted.

I don't get it, what is bad about what they are doing? People need a CDN, they choose the one they find the best.

The discussion revolves around the equivalent of taking nutrition advice from Coca-Cola's blog.

escalated quickly

Re: Cloudflare CEO is lying to you about the bot traffic jump

#83

I deal with scrapers that sometimes border on DDoSes for LessWrong. The amount of bot traffic varies greatly between sites; if you have more URLs you get more bot traffic (regardless of whether those URLs represent a deep content catalog, or useless URL parameter permutations). It's bad for LW because of the content-catalog depth. It's easy to drastically underestimate the amount of bot traffic, because bots make eff…

Meta comes through with a /24 worth of scrapers and ignores robots.txt. I'm inclined to poison my data with fake information about Zuckerberg.

Re: Cloudflare CEO is lying to you about the bot traffic jump

#84
post #60

Earlier quoted context omitted.

> CF serves a customers need CF serves something it convinced customers they need. Static blogs hiding behind bot protection (in some cases blocking legit users from GrapheneOS because it's difficult to fingerprint them) because someone convinced them they'll be DDoSed by bots otherwise is a loss to the Internet. A lot of self-hosters running CF tunnels because they don't know better also contributes. > It's more hea…

[flagged]

It's not about being incompetent. Quite often in selfhosted subreddits and forums you will see people surprised that Cloudflare can see their traffic in plaintext.

Of course, they probably don't, but the fact that they can and that their policies now influence XX% of internet traffic is bad for the open internet.

Re: Cloudflare CEO is lying to you about the bot traffic jump

#86
post #83

I deal with scrapers that sometimes border on DDoSes for LessWrong. The amount of bot traffic varies greatly between sites; if you have more URLs you get more bot traffic (regardless of whether those URLs represent a deep content catalog, or useless URL parameter permutations). It's bad for LW because of the content-catalog depth. It's easy to drastically underestimate the amount of bot traffic, because bots make eff…

Meta comes through with a /24 worth of scrapers and ignores robots.txt. I'm inclined to poison my data with fake information about Zuckerberg.

Did you check IP addresses, are they all from AS32934?

Re: Cloudflare CEO is lying to you about the bot traffic jump

#87
post #75
post #60

Earlier quoted context omitted.

> CF serves a customers need CF serves something it convinced customers they need. Static blogs hiding behind bot protection (in some cases blocking legit users from GrapheneOS because it's difficult to fingerprint them) because someone convinced them they'll be DDoSed by bots otherwise is a loss to the Internet. A lot of self-hosters running CF tunnels because they don't know better also contributes. > It's more hea…

> A lot of self-hosters running CF tunnels because they don't know better also contributes. Are you saying CF documentation is better than Computer Science / Networking education resources? Why don't people know better? I thought the tunnels are mostly used to bypass NAT's. > Static blogs hiding behind bot protection I'm not sure what is the proportion of the static vs dynamic sites, but I would argue that for wordpr…

> I thought the tunnels are mostly used to bypass NAT's.

While not free, you can do with with TCP HAProxy streams on a cheap VPS. A lot of people using them to bypass NAT don't realise that Cloudflare decrypt the traffic on the way - that's what I meant about them not knowing better.

Re: Cloudflare CEO is lying to you about the bot traffic jump

#88
post #58
post #34

Earlier quoted context omitted.

(I'm ex CF) This is backwards. Nobody "allowed" anyting. CF serves a customers need. You can argue with the solution but you can't argue with the core problem. It's more healthy to start the conversation of _why_ CF services are valuable.

Regulation could conceivably disallow a single company to control such a significant portion of internet traffic. The parent can be interpreted as lamenting the absence of such regulation.

The only thing more annoying than people chanting "regulation is bad" is people chanting "regulation is good".

What regulation? Be specific.

CloudFlare provides significant utility to me. I chose to use them. Explain why you think someone else needs to butt into this relationship.

Re: Cloudflare CEO is lying to you about the bot traffic jump

#89
post #11

"Cloudflare CEO is lying" is a bit of an aggressive take when he linked to the exact data so you can see it for yourself - and that's how this article was able to analyze it: https://radar.cloudflare.com/traffic#bot-vs-human Update: I see the problem. Here's the full tweet: https://x.com/eastdakota/status/2062212701414187452 "Thought it would be end of 2027, then early 2027, but agentic traffic growing so fast that b…

The tweet implied there was a shift but for some reason radar is cutting off everything except recent data for just this one time series so we can't verify his storyline.

Re: Cloudflare CEO is lying to you about the bot traffic jump

#90
post #83

Earlier quoted context omitted.

Meta comes through with a /24 worth of scrapers and ignores robots.txt. I'm inclined to poison my data with fake information about Zuckerberg.

Did you check IP addresses, are they all from AS32934?

Yes

57.141.0.42 - - [05/Jun/2026:19:50:19 +0000] "GET /mid/a017bc62-0982-42db-8403-241d69da8d0f@alexander-goetzenstein.my-fqdn.de HTTP/2.0" 303 0 "-" "meta-externalagent/1.1 (+https://developers.facebook.com/docs/sharing/webmasters/craw...)"

57.141.0.48 - - [05/Jun/2026:19:50:22 +0000] "GET /group/comp.os.linux.advocacy/a/a236f5a5-63a4-4982-8bb6-07ffc684201b@googlegroups.com HTTP/2.0" 200 34838 "-" "meta-externalagent/1.1 (+https://developers.facebook.com/docs/sharing/webmasters/craw...)"

57.141.0.55 - - [05/Jun/2026:19:50:23 +0000] "GET /group/alt.recovery.aa/a/ne6onq%24hpp%241@dont-email.me HTTP/2.0" 200 5606 "-" "meta-externalagent/1.1 (+https://developers.facebook.com/docs/sharing/webmasters/craw...)"

57.141.0.56 - - [05/Jun/2026:19:50:24 +0000] "GET /group/aioe.news.assistenza/a/qpukie%241i1g%241@neodome.net?view=headers HTTP/2.0" 200 17027 "-" "meta-externalagent/1.1 (+https://developers.facebook.com/docs/sharing/webmasters/craw...)"

57.141.0.36 - - [05/Jun/2026:19:50:29 +0000] "GET /group/alt.obituaries/a/uf8pej%241hqi1%241@news.xmission.com HTTP/2.0" 200 6123 "-" "meta-externalagent/1.1 (+https://developers.facebook.com/docs/sharing/webmasters/craw...)"

57.141.0.66 - - [05/Jun/2026:19:50:29 +0000] "GET /group/comp.theory/a/v3640k%24vg63%243@dont-email.me HTTP/2.0" 200 148720 "-" "meta-externalagent/1.1 (+https://developers.facebook.com/docs/sharing/webmasters/craw...)"

Post reply on HN