I deal with scrapers that sometimes border on DDoSes for LessWrong. The amount of bot traffic varies greatly between sites; if you have more URLs you get more bot traffic (regardless of whether those URLs represent a deep content catalog, or useless URL parameter permutations). It's bad for LW because of the content-catalog depth. It's easy to drastically underestimate the amount of bot traffic, because bots make eff…
Cloudflare CEO is lying to you about the bot traffic jump
81–90 of 149 posts
Re: Cloudflare CEO is lying to you about the bot traffic jump
#82Re: Cloudflare CEO is lying to you about the bot traffic jump
#83I deal with scrapers that sometimes border on DDoSes for LessWrong. The amount of bot traffic varies greatly between sites; if you have more URLs you get more bot traffic (regardless of whether those URLs represent a deep content catalog, or useless URL parameter permutations). It's bad for LW because of the content-catalog depth. It's easy to drastically underestimate the amount of bot traffic, because bots make eff…
Re: Cloudflare CEO is lying to you about the bot traffic jump
#84Earlier quoted context omitted.
> CF serves a customers need CF serves something it convinced customers they need. Static blogs hiding behind bot protection (in some cases blocking legit users from GrapheneOS because it's difficult to fingerprint them) because someone convinced them they'll be DDoSed by bots otherwise is a loss to the Internet. A lot of self-hosters running CF tunnels because they don't know better also contributes. > It's more hea…
[flagged]
Of course, they probably don't, but the fact that they can and that their policies now influence XX% of internet traffic is bad for the open internet.
Re: Cloudflare CEO is lying to you about the bot traffic jump
#85Re: Cloudflare CEO is lying to you about the bot traffic jump
#86I deal with scrapers that sometimes border on DDoSes for LessWrong. The amount of bot traffic varies greatly between sites; if you have more URLs you get more bot traffic (regardless of whether those URLs represent a deep content catalog, or useless URL parameter permutations). It's bad for LW because of the content-catalog depth. It's easy to drastically underestimate the amount of bot traffic, because bots make eff…
Meta comes through with a /24 worth of scrapers and ignores robots.txt. I'm inclined to poison my data with fake information about Zuckerberg.
Re: Cloudflare CEO is lying to you about the bot traffic jump
#87Earlier quoted context omitted.
> CF serves a customers need CF serves something it convinced customers they need. Static blogs hiding behind bot protection (in some cases blocking legit users from GrapheneOS because it's difficult to fingerprint them) because someone convinced them they'll be DDoSed by bots otherwise is a loss to the Internet. A lot of self-hosters running CF tunnels because they don't know better also contributes. > It's more hea…
> A lot of self-hosters running CF tunnels because they don't know better also contributes. Are you saying CF documentation is better than Computer Science / Networking education resources? Why don't people know better? I thought the tunnels are mostly used to bypass NAT's. > Static blogs hiding behind bot protection I'm not sure what is the proportion of the static vs dynamic sites, but I would argue that for wordpr…
While not free, you can do with with TCP HAProxy streams on a cheap VPS. A lot of people using them to bypass NAT don't realise that Cloudflare decrypt the traffic on the way - that's what I meant about them not knowing better.
Re: Cloudflare CEO is lying to you about the bot traffic jump
#88Earlier quoted context omitted.
(I'm ex CF) This is backwards. Nobody "allowed" anyting. CF serves a customers need. You can argue with the solution but you can't argue with the core problem. It's more healthy to start the conversation of _why_ CF services are valuable.
Regulation could conceivably disallow a single company to control such a significant portion of internet traffic. The parent can be interpreted as lamenting the absence of such regulation.
What regulation? Be specific.
CloudFlare provides significant utility to me. I chose to use them. Explain why you think someone else needs to butt into this relationship.
Re: Cloudflare CEO is lying to you about the bot traffic jump
#89"Cloudflare CEO is lying" is a bit of an aggressive take when he linked to the exact data so you can see it for yourself - and that's how this article was able to analyze it: https://radar.cloudflare.com/traffic#bot-vs-human Update: I see the problem. Here's the full tweet: https://x.com/eastdakota/status/2062212701414187452 "Thought it would be end of 2027, then early 2027, but agentic traffic growing so fast that b…
Re: Cloudflare CEO is lying to you about the bot traffic jump
#90Earlier quoted context omitted.
Meta comes through with a /24 worth of scrapers and ignores robots.txt. I'm inclined to poison my data with fake information about Zuckerberg.
Did you check IP addresses, are they all from AS32934?
57.141.0.42 - - [05/Jun/2026:19:50:19 +0000] "GET /mid/a017bc62-0982-42db-8403-241d69da8d0f@alexander-goetzenstein.my-fqdn.de HTTP/2.0" 303 0 "-" "meta-externalagent/1.1 (+https://developers.facebook.com/docs/sharing/webmasters/craw...)"
57.141.0.48 - - [05/Jun/2026:19:50:22 +0000] "GET /group/comp.os.linux.advocacy/a/a236f5a5-63a4-4982-8bb6-07ffc684201b@googlegroups.com HTTP/2.0" 200 34838 "-" "meta-externalagent/1.1 (+https://developers.facebook.com/docs/sharing/webmasters/craw...)"
57.141.0.55 - - [05/Jun/2026:19:50:23 +0000] "GET /group/alt.recovery.aa/a/ne6onq%24hpp%241@dont-email.me HTTP/2.0" 200 5606 "-" "meta-externalagent/1.1 (+https://developers.facebook.com/docs/sharing/webmasters/craw...)"
57.141.0.56 - - [05/Jun/2026:19:50:24 +0000] "GET /group/aioe.news.assistenza/a/qpukie%241i1g%241@neodome.net?view=headers HTTP/2.0" 200 17027 "-" "meta-externalagent/1.1 (+https://developers.facebook.com/docs/sharing/webmasters/craw...)"
57.141.0.36 - - [05/Jun/2026:19:50:29 +0000] "GET /group/alt.obituaries/a/uf8pej%241hqi1%241@news.xmission.com HTTP/2.0" 200 6123 "-" "meta-externalagent/1.1 (+https://developers.facebook.com/docs/sharing/webmasters/craw...)"
57.141.0.66 - - [05/Jun/2026:19:50:29 +0000] "GET /group/comp.theory/a/v3640k%24vg63%243@dont-email.me HTTP/2.0" 200 148720 "-" "meta-externalagent/1.1 (+https://developers.facebook.com/docs/sharing/webmasters/craw...)"