Live data from Hacker News

Post Mortem on Cloudflare Control Plane and Analytics Outage

blog.cloudflare.com

61–70 of 241 posts

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#61
post #50

Earlier quoted context omitted.

And that this was unironically written in the same post mortem: “We are good at distributed systems.” There’s a lack of awareness there.

Well, they did distribute their systems. Some were in the running DC, some were not ;)

They are good at systems that are distributed; they are very bad at ensuring systems they sell thier custoners are distributed.

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#62
post #36

A lot of mud-slinging on here about HA setup and CF's dealing with the problem but I can only assume people are armchair experts with no real experience of HA at the scale of CF. "So the root cause for the outage was that they relied on a single data center.". No. Root cause was that data centre operator didn't manage the outage properly and didn't have systems in place in which case they could have avoided it + some…

Couldn’t agree more.

Many of these comments sound like they’re coming from some mythical alternate universe where bugs don’t exist and people and orgs have 100% flawless execution every time.

It reminds me a little of someone sitting at a sports bar yelling about a “stupid” play or otherwise criticizing a 0.0001% athlete who is playing at a level they can’t possibly fathom.

Monday Morning quarterbacking.

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#63

hot take: HN is way too biased and sympathetic towards the provider whenever an outage like this happens.

Not a lot of measured takes here. It seems to be either “eh we get it, comms could have been better” or “they’re idiots”.

The first group of people have been to war. The second have not.

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#64
post #11
post #5

Cloudflare's control plane and analytics systems run primarily on servers in three data centers around Hillsboro, Oregon. The three data centers are independent of one another, each have multiple utility power feeds, and each have multiple redundant and independent network connections. The facilities were intentionally chosen to be at a distance apart that would minimize the chances that a natural disaster would caus…

> Hillsboro, Oregon, an earthquake could probably take out all three simultaneously. Is it west of I5? (yes) Oh yeah, they all gone. Cascadia Subduction Zone - https://pnsn.org/outreach/earthquakesources/csz

Wikipedia's entry for Hillsboro: "Elevation 194 ft (60 m)"

Between that, and being ~50 miles inland - I'd say there's ~zero threat of Cascadia quakes or tsunamis directly knocking out those DC's. (Yeah, larger-scale infrastructure and social order could still be killers.)

OTOH - Mt. St. Helens is about 60 miles NNE of Hillsboro. If that really went boom, and the wind was right...how many cm's of dry volcanic ash can the roofs of those DC's bear? What if rain wets that ash? How about their HVAC systems' filters?

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#65
post #54

As someone who was slightly affected by this outage, I personally also find this post-mortem to be lacking. 75% of the post-mortem talks about the power outage at PDX-04 and blames Flexential. Okay, fair - it was a bit of a disaster what was happening there judging from the text. But by end of November 2 (UTC), power was fully restored. It still took ~30 hours according to the post-mortem for Cloudflare to fully reco…

> Also one aspect I am missing here is the lack of communication - especially to Enterprise customers.

They blame Flexential for lack of communication, but were the first one not saying anything.

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#66
> We are a relatively large customer of the facility, consuming approximately 10 percent of its total capacity.

I'm surprised that CF are renting space in colocation facilities. I would have expected a business of their size to have their own DCs. Is this common practice for cloud providers?

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#67
post #11
post #5

Cloudflare's control plane and analytics systems run primarily on servers in three data centers around Hillsboro, Oregon. The three data centers are independent of one another, each have multiple utility power feeds, and each have multiple redundant and independent network connections. The facilities were intentionally chosen to be at a distance apart that would minimize the chances that a natural disaster would caus…

> Hillsboro, Oregon, an earthquake could probably take out all three simultaneously. Is it west of I5? (yes) Oh yeah, they all gone. Cascadia Subduction Zone - https://pnsn.org/outreach/earthquakesources/csz

And most of thier SREs. Spending 30 hours to recover from the worst natural disaster in recorded history is slightly diffrent then from a ground fault on a single transformer.

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#68
post #6

I think overall Cloudflare did a decent job on this. Clearly the DC provider cocked up big time here, but Cloudflare kept running fine for the vast majority of customers globally. No system is perfect and it’s only apocalyptic scenarios like this where the vulnerabilities are exposed - and they will now be fixed. Hope the SRE guys got some rest after all that stress.

> Clearly the DC provider cocked up big time here

Actually, this is the CF version, maybe Flexential will come out with a different one.

BTW, if you design a system to survive a DC failure, you cannot blame the DC failure.

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#69

> We are a relatively large customer of the facility, consuming approximately 10 percent of its total capacity. I'm surprised that CF are renting space in colocation facilities. I would have expected a business of their size to have their own DCs. Is this common practice for cloud providers?

You thought that they build > 300 DC's?

Colo is much more flexible, cheaper and quicker to start. Definitely since they sit close to the end-user on the data plane.

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#70
post #58

Earlier quoted context omitted.

From what I was reading on the status page & customers here on HN, WARP + Zero Trust were also majorly affected, which would be quite impactful for a company using these products for their internal authentication. It's not just streams, image upload & Logpush.

Those customers were impacted until the DC was back up ( 1-2 hours?) On the config plane. The data plane ( which I mentioned) had no issues. It's literally in the title what was affected: "Post Mortem on Cloudflare Control Plane and Analytics Outage" Eg. The status page mentioned the healthchecks not working, while everything was fine with it. There were just no analytics at that time to confirm that. Source: I watch…

> Those customers were impacted until the DC was back up ( 1-2 hours?) On the config plane.

Which was still like ~12+ hours, if we check the status page.

>Eg. The status page mentioned the healthchecks not working, while everything was fine with it. There were just no analytics at that time to confirm that.

What good is a status page that's lying to you? Especially since CF manually updates it, anyway?

>Source: I watched it all happen in the cloudflare discord channel.

Wow, as a business customer I definitely like watching some Discord channel for status updates.

Post reply on HN