Earlier quoted context omitted.
And that this was unironically written in the same post mortem: “We are good at distributed systems.” There’s a lack of awareness there.
Well, they did distribute their systems. Some were in the running DC, some were not ;)
Post Mortem on Cloudflare Control Plane and Analytics Outage
61–70 of 241 posts
Re: Post Mortem on Cloudflare Control Plane and Analytics Outage
#62A lot of mud-slinging on here about HA setup and CF's dealing with the problem but I can only assume people are armchair experts with no real experience of HA at the scale of CF. "So the root cause for the outage was that they relied on a single data center.". No. Root cause was that data centre operator didn't manage the outage properly and didn't have systems in place in which case they could have avoided it + some…
Many of these comments sound like they’re coming from some mythical alternate universe where bugs don’t exist and people and orgs have 100% flawless execution every time.
It reminds me a little of someone sitting at a sports bar yelling about a “stupid” play or otherwise criticizing a 0.0001% athlete who is playing at a level they can’t possibly fathom.
Monday Morning quarterbacking.
Re: Post Mortem on Cloudflare Control Plane and Analytics Outage
#63hot take: HN is way too biased and sympathetic towards the provider whenever an outage like this happens.
The first group of people have been to war. The second have not.
Re: Post Mortem on Cloudflare Control Plane and Analytics Outage
#64Cloudflare's control plane and analytics systems run primarily on servers in three data centers around Hillsboro, Oregon. The three data centers are independent of one another, each have multiple utility power feeds, and each have multiple redundant and independent network connections. The facilities were intentionally chosen to be at a distance apart that would minimize the chances that a natural disaster would caus…
> Hillsboro, Oregon, an earthquake could probably take out all three simultaneously. Is it west of I5? (yes) Oh yeah, they all gone. Cascadia Subduction Zone - https://pnsn.org/outreach/earthquakesources/csz
Between that, and being ~50 miles inland - I'd say there's ~zero threat of Cascadia quakes or tsunamis directly knocking out those DC's. (Yeah, larger-scale infrastructure and social order could still be killers.)
OTOH - Mt. St. Helens is about 60 miles NNE of Hillsboro. If that really went boom, and the wind was right...how many cm's of dry volcanic ash can the roofs of those DC's bear? What if rain wets that ash? How about their HVAC systems' filters?
Re: Post Mortem on Cloudflare Control Plane and Analytics Outage
#65As someone who was slightly affected by this outage, I personally also find this post-mortem to be lacking. 75% of the post-mortem talks about the power outage at PDX-04 and blames Flexential. Okay, fair - it was a bit of a disaster what was happening there judging from the text. But by end of November 2 (UTC), power was fully restored. It still took ~30 hours according to the post-mortem for Cloudflare to fully reco…
They blame Flexential for lack of communication, but were the first one not saying anything.
Re: Post Mortem on Cloudflare Control Plane and Analytics Outage
#66I'm surprised that CF are renting space in colocation facilities. I would have expected a business of their size to have their own DCs. Is this common practice for cloud providers?
Re: Post Mortem on Cloudflare Control Plane and Analytics Outage
#67Cloudflare's control plane and analytics systems run primarily on servers in three data centers around Hillsboro, Oregon. The three data centers are independent of one another, each have multiple utility power feeds, and each have multiple redundant and independent network connections. The facilities were intentionally chosen to be at a distance apart that would minimize the chances that a natural disaster would caus…
> Hillsboro, Oregon, an earthquake could probably take out all three simultaneously. Is it west of I5? (yes) Oh yeah, they all gone. Cascadia Subduction Zone - https://pnsn.org/outreach/earthquakesources/csz
Re: Post Mortem on Cloudflare Control Plane and Analytics Outage
#68I think overall Cloudflare did a decent job on this. Clearly the DC provider cocked up big time here, but Cloudflare kept running fine for the vast majority of customers globally. No system is perfect and it’s only apocalyptic scenarios like this where the vulnerabilities are exposed - and they will now be fixed. Hope the SRE guys got some rest after all that stress.
Actually, this is the CF version, maybe Flexential will come out with a different one.
BTW, if you design a system to survive a DC failure, you cannot blame the DC failure.
Re: Post Mortem on Cloudflare Control Plane and Analytics Outage
#69> We are a relatively large customer of the facility, consuming approximately 10 percent of its total capacity. I'm surprised that CF are renting space in colocation facilities. I would have expected a business of their size to have their own DCs. Is this common practice for cloud providers?
Colo is much more flexible, cheaper and quicker to start. Definitely since they sit close to the end-user on the data plane.
Re: Post Mortem on Cloudflare Control Plane and Analytics Outage
#70Earlier quoted context omitted.
From what I was reading on the status page & customers here on HN, WARP + Zero Trust were also majorly affected, which would be quite impactful for a company using these products for their internal authentication. It's not just streams, image upload & Logpush.
Those customers were impacted until the DC was back up ( 1-2 hours?) On the config plane. The data plane ( which I mentioned) had no issues. It's literally in the title what was affected: "Post Mortem on Cloudflare Control Plane and Analytics Outage" Eg. The status page mentioned the healthchecks not working, while everything was fine with it. There were just no analytics at that time to confirm that. Source: I watch…
Which was still like ~12+ hours, if we check the status page.
>Eg. The status page mentioned the healthchecks not working, while everything was fine with it. There were just no analytics at that time to confirm that.
What good is a status page that's lying to you? Especially since CF manually updates it, anyway?
>Source: I watched it all happen in the cloudflare discord channel.
Wow, as a business customer I definitely like watching some Discord channel for status updates.