Live data from Hacker News

Cloudflare incident on August 21, 2025

blog.cloudflare.com

21–30 of 46 posts

Re: Cloudflare incident on August 21, 2025

#22
I'm having trouble understanding the second diagram in the article. I can make sense of a directed graph, but this one has thin horizontal lines with arrows leaving them in both directions. These lines look like dividers, not nodes, so I'm not sure how to interpret it.

Re: Cloudflare incident on August 21, 2025

#24
> The incident was a result of a surge of traffic from a single customer that overloaded Cloudflare's links with AWS us-east-1. It was a network congestion event, not an attack or a BGP hijack.

And no one knew a single thing about it until the incident. That is the current network management state of the art, let Cloudflare deal.

Re: Cloudflare incident on August 21, 2025

#25
post #23

Only real long term mitigation is to move to another aws region; us-east-1 seems to suffer from all kinds of scaling challenges.

There's nothing to suggest the link between Cloudflare and any other AWS region has more capacity or that there aren't more disruptive Cloudflare customers using those regions.

Re: Cloudflare incident on August 21, 2025

#26
post #7
post #2

Wild that one tenant’s cache-hit traffic could tip over Cloudflare’s interconnect capacity

You'd be surprised how low the capacity of a lot of internet links is. 10Gbps is common on smaller networks - let me rephrase that, a small to medium ISP might only have 10Gbps to each of most of their peering partners. Normally, traffic is distributed, going to different places, coming from different places, and each link is partially utilized. But unusual patterns can fill up one specific link. 10Gbps is old techno…

Future proofing inevitable things should be something to talk about more.

For instance, people will be scraping at a "growing" rate as they figure out how everything AI works. We might as well figure out some standard seeded data packages for training that ~all sources/sectors agree to make available as public torrents to reduce this type of problem.

[I realize this ask is currently idealistic, but it's an anchor point to negotiate from.]

Re: Cloudflare incident on August 21, 2025

#27
post #2

Wild that one tenant’s cache-hit traffic could tip over Cloudflare’s interconnect capacity

That's what started the incident. It was prolonged by the fact that Cloudflare didn't react correctly to withdrawn BGP routes to a major peer, that the secondary routes had reduced capacity due to unaddressed problems, and basic nuisance rate limiting had to be done manually. It seems like they just build huge peering pipes and basically just hope for the best. They've maybe gotten so used to this working that they'l…

Wasn’t the problem exacerbated precisely by withdrawing a BGP link because all the same traffic is then forced over a smaller number of physical links?

Re: Cloudflare incident on August 21, 2025

#28
post #23

Only real long term mitigation is to move to another aws region; us-east-1 seems to suffer from all kinds of scaling challenges.

There's nothing to suggest the link between Cloudflare and any other AWS region has more capacity or that there aren't more disruptive Cloudflare customers using those regions.

But there is absolutely something to suggest "if you only support one region for some tasks, you're going to have problems that other people don't have."

Re: Cloudflare incident on August 21, 2025

#30
post #8

Anyone want to tell Cloudflare that BGP advertisements at AWS are automated and their congested network directly cause BGP withdrawals as the automated system detected congestion and decreased traffic to remediate it?

The way I read the blog post, it seems they're very aware of that.

I imagine Cloudflare and AWS were on a Chime bridge while this all went down, they both have a lot at stake here.

Post reply on HN