Live data from Hacker News

Post Mortem on Cloudflare Control Plane and Analytics Outage

blog.cloudflare.com

91–100 of 241 posts

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#91
post #12

Not criticism, just remarks: > While, over time, our practice is to migrate the backend for these services to our best practices, we did not formally require that before products were declared generally available (GA). I really like the model where a single team in a company, with Product + Dev, can quickly ship, iterate on a new product, and prove market demand without going through layers and layers of internal bur…

> Security

Security is what keeps a single service getting breached from causing the whole company to get breached.

> Privacy/Legal

Cloudflare doesn't get indemnification from the law just because a customer agrees to mutually break the law.

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#92
post #84

They don't always know it, but all large systems are moving gradually towards dependency management system with logic rules that covers "everything", physical, logical, human and administrative dependencies. Every time something new not covered is discovered, new rules and conditions are added. You can do it with manual checklists, multiple rule checkers, or put everything together. I suspect that in end it's just ea…

This is such an interesting way of putting it. I think this has been the subconscious reason I've been gravitating towards defining _everything_ I manage personally (and not yet at work) with Nix. It's not quite to the extent you're talking about here, of course, but in a similar vein at least.

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#93
post #11

Earlier quoted context omitted.

> Hillsboro, Oregon, an earthquake could probably take out all three simultaneously. Is it west of I5? (yes) Oh yeah, they all gone. Cascadia Subduction Zone - https://pnsn.org/outreach/earthquakesources/csz

Wikipedia's entry for Hillsboro: "Elevation 194 ft (60 m)" Between that, and being ~50 miles inland - I'd say there's ~zero threat of Cascadia quakes or tsunamis directly knocking out those DC's. (Yeah, larger-scale infrastructure and social order could still be killers.) OTOH - Mt. St. Helens is about 60 miles NNE of Hillsboro. If that really went boom, and the wind was right...how many cm's of dry volcanic ash can…

I was never worried about the tsunami. Okay, maybe not gone, but I wouldn't say it would be operational.

https://www.oregon.gov/oem/Documents/Cascadia_Rising_Exercis...

50% of roads and near 75% of bridges damaged on the west coast and the I5 corridor.

Refer to PDF page #93 where over 70% of power generation is highly damaged on the I5 corridor and 60% in the coastal areas with 0% undamaged.

Highly damaged - "Extensive damage to generation plants, substations, and buildings. Repairs are needed to regain functionality. Restoring power to meet 90% of demand may take months to one year."

"In the immediate aftermath of the earthquake, cities within 100 miles of the Pacific coastline may experience partial or complete blackout. Seventy percent of the electric facilities in the I-5 corridor may suffer considerable damage to generation plants, and many distribution circuits and substations may fail, resulting in a loss of over half of the systems load capacity (see Table 22). Most electrical power assets on the coast may suffer damage severe enough as to render the equipment and structures irreparable"

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#94

> However, we had never tested fully taking the entire PDX-04 facility offline. That is a painful lesson, but unless you are physically powering off the dc or at least disconnecting the network from the outside world you are not testing a real disaster. You can point fingers at the facility operators, but at the end of the day you have to be able to recover from a dc going completely offline and maybe never coming ba…

This is a fair point. Imagine there had been a serious fire like OVH suffered or flooding that destroyed the data center. Would Cloudflare have been able to recover?

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#95
post #48
post #21

> While most of our critical control plane systems had been migrated to the high availability cluster, some services, especially for some newer products, had not yet been added to the high availability cluster. > The other two data centers running in the area would take over responsibility for the high availability cluster and keep critical services online. Generally that worked as planned. Unfortunately, we discover…

> I find that pretty embarrassing for a company like Cloudflare, which powers such relevant parts of the internet. Bah, who cares about such unimportant details, what's important is that ~dev velocity~ was reaaally high right until that moment! > We were also far too lax about requiring new products and their associated databases to integrate with the high availability cluster. Cloudflare allows multiple teams to inn…

I’m going to leave out some details but there was a period of time where you could bypass cloudflare’s IP whitelisting by using Apple’s iCloud relay service. This was fixed but to my knowledge never disclosed.

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#96

Earlier quoted context omitted.

You thought that they build > 300 DC's? Colo is much more flexible, cheaper and quicker to start. Definitely since they sit close to the end-user on the data plane.

> You thought that they build > 300 DC's? I have no idea how many DCs they have or operate in. Where does "300" come from? > Colo is much more flexible, cheaper and quicker to start. Definitely since they sit close to the end-user on the data plane. I understand that, but it has the disadvantage of reduced control and observability - particularly in the event of an outage such as that described in the blog post. I ki…

https://www.cloudflare.com/network/

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#97
post #70

Earlier quoted context omitted.

Those customers were impacted until the DC was back up ( 1-2 hours?) On the config plane. The data plane ( which I mentioned) had no issues. It's literally in the title what was affected: "Post Mortem on Cloudflare Control Plane and Analytics Outage" Eg. The status page mentioned the healthchecks not working, while everything was fine with it. There were just no analytics at that time to confirm that. Source: I watch…

> Those customers were impacted until the DC was back up ( 1-2 hours?) On the config plane. Which was still like ~12+ hours, if we check the status page. >Eg. The status page mentioned the healthchecks not working, while everything was fine with it. There were just no analytics at that time to confirm that. What good is a status page that's lying to you? Especially since CF manually updates it, anyway? >Source: I wat…

?

This wasn't about status updates going to discord only.

There is literally a discussion section on the discord, named: #general-discussions

Not everything was clear in the discord too ( eg. The healthchecks were discussed there), that's not something you want to copy-paste in the status updates...

Priority for cloudflare seemed to get everything back up. And what they thought was down, was always mentioned in the status updates.

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#98
post #65

Earlier quoted context omitted.

> Also one aspect I am missing here is the lack of communication - especially to Enterprise customers. They blame Flexential for lack of communication, but were the first one not saying anything.

Even "we don't know why our data center is failing, but we're sending a team over to physically investigate now" would have been A+ communication in the moment.

Everything was on the status page since the start?

DC related updates:

> Update - Power to Cloudflare’s core North America data center has been partially restored. Cloudflare has failed over some core services to a backup data center, which has partially remediated impact. Cloudflare is currently working to restore the remaining affected services and bring the core North America data center back online. Nov 02, 2023 - 17:08 UTC

> Identified - Cloudflare is assessing a loss of power impacting data centres while simultaneously failing over services.

We will keep providing regular updates until the issue is resolved, thank you for your patience as we work on mitigating the problem. Nov 02, 2023 - 13:40 UTC

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#99

Earlier quoted context omitted.

You thought that they build > 300 DC's? Colo is much more flexible, cheaper and quicker to start. Definitely since they sit close to the end-user on the data plane.

> You thought that they build > 300 DC's? I have no idea how many DCs they have or operate in. Where does "300" come from? > Colo is much more flexible, cheaper and quicker to start. Definitely since they sit close to the end-user on the data plane. I understand that, but it has the disadvantage of reduced control and observability - particularly in the event of an outage such as that described in the blog post. I ki…

Probably most of us follow Cloudflare a bit more closely.

They want DC's close to every big city. I think most of us knew that they can't launch > 300 DC's in such a short amount of time.

The many amount of DC's is mentioned a lot ( social networks, blogs, here).

There is a distinction between eg. AWS / Azure / ... Which work with a couple of big DC's, while cloudflare operates more spread across more locations.

You're comment did made me realize it may may not be that clear from an outsider viewpoint though ( fyi, I'm an outsider too)

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#100

Interesting choice to spend the bulk of the article publicly shifting blame to a vendor by name and speculating on their root cause. Also an interesting choice to publicly call out that you're a whale in the facility and include an electrical diagram clearly marked Confidential by your vendor in the postmortem. Honestly, this is rather unprofessional. I understand and support explaining what triggered the event and g…

Especially since it shouldn't matter why the DC failed — Cloudflare's entire business model is selling services allegedly designed to survive that. 99% of the fault lies with Cloudflare for not being able to do their core job.
Post reply on HN