Live data from Hacker News

Post Mortem on Cloudflare Control Plane and Analytics Outage

blog.cloudflare.com

101–110 of 241 posts

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#101

> However, we had never tested fully taking the entire PDX-04 facility offline. That is a painful lesson, but unless you are physically powering off the dc or at least disconnecting the network from the outside world you are not testing a real disaster. You can point fingers at the facility operators, but at the end of the day you have to be able to recover from a dc going completely offline and maybe never coming ba…

This is a fair point. Imagine there had been a serious fire like OVH suffered or flooding that destroyed the data center. Would Cloudflare have been able to recover?

Most likely, yes. They have enough customer lock-in that enough customers would stick with them even if it took them a week to rebuild everything from in other DCs.

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#102

Interesting choice to spend the bulk of the article publicly shifting blame to a vendor by name and speculating on their root cause. Also an interesting choice to publicly call out that you're a whale in the facility and include an electrical diagram clearly marked Confidential by your vendor in the postmortem. Honestly, this is rather unprofessional. I understand and support explaining what triggered the event and g…

[deleted]

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#103
I asked a sales rep once about services going out and how that would affect CF For Teams. They said it would be virtually impossible for CF to go down because of all their data centers around the world. Paraphrasing, “if there’s an outage, there’s definitely something going wrong with the internet.”

And here we are. My trust in them has hit zero.

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#104

I asked a sales rep once about services going out and how that would affect CF For Teams. They said it would be virtually impossible for CF to go down because of all their data centers around the world. Paraphrasing, “if there’s an outage, there’s definitely something going wrong with the internet.” And here we are. My trust in them has hit zero.

Why would you trust a sales rep?

Even honest engineers cannot foresee the exact cascading consequences effects of such outages. Sales reps are not paid to be either competent on such issues nor to be honest.

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#105

I asked a sales rep once about services going out and how that would affect CF For Teams. They said it would be virtually impossible for CF to go down because of all their data centers around the world. Paraphrasing, “if there’s an outage, there’s definitely something going wrong with the internet.” And here we are. My trust in them has hit zero.

FWIW, I'm a Cloudflare Enterprise customer and we had zero downtime. Only thing that was temporarily unavailable was the cloudflare dashboard.

I feel like a lot of people in this thread are commenting under the impression that all of Cloudflare was down for 24 hours when in reality I wouldn't be surprised if a lot of customers were unaffected and unaware of the incident.

I wouldn't even have known of the outage had it not been for HN..

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#106

Interesting choice to spend the bulk of the article publicly shifting blame to a vendor by name and speculating on their root cause. Also an interesting choice to publicly call out that you're a whale in the facility and include an electrical diagram clearly marked Confidential by your vendor in the postmortem. Honestly, this is rather unprofessional. I understand and support explaining what triggered the event and g…

Especially since it shouldn't matter why the DC failed — Cloudflare's entire business model is selling services allegedly designed to survive that. 99% of the fault lies with Cloudflare for not being able to do their core job.

In all fairness the rest of the article is about that

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#107

I love how thorough Cloudflare post mortem’s are. Reading the frank, transparent explanations are like a breath of fresh air compared to the obfuscation of nearly every other company comm’s strategy. We were affected but it’s blog posts like these that make me never want to move away. Everyone makes mistakes. Everyone has bad days. It’s how you react afterwards that makes the difference.

I would generally agree with you, but this post mortem was 75% blaming Flexential even though it took them almost two days to recover after power was restored. The power outage should have been a single paragraph and then pivoted - DC failures happen, its part of life. Failing to properly account for and recover from it is where the real learnings for Cloudflare are.

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#108

I asked a sales rep once about services going out and how that would affect CF For Teams. They said it would be virtually impossible for CF to go down because of all their data centers around the world. Paraphrasing, “if there’s an outage, there’s definitely something going wrong with the internet.” And here we are. My trust in them has hit zero.

Trust? Random sample from the last 60 days...

Cloudflare outage – 24 hours now - https://news.ycombinator.com/item?id=38112515

Cloudflare Dashboard Logins Failing - https://news.ycombinator.com/item?id=38112230

Ask HN: Cloudflare Workers are down? - https://news.ycombinator.com/item?id=38074906

Cloudflare API, dashboard, tunnels down - https://news.ycombinator.com/item?id=38014582

Cloudflare Intermittent API Failures for Cloudflare Pages, Workers and Images - https://news.ycombinator.com/item?id=37819045

Cloudflare Issues with 1.1.1.1 public resolver and WARP - https://news.ycombinator.com/item?id=37762731

Cloudflare – Network Performance Issues - https://news.ycombinator.com/item?id=37604609

Cloudflare Issues Passing Challenge Pages - https://news.ycombinator.com/item?id=37336743

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#109
post #55
post #34

They really threw the electricity power provider under the bus there.

The electricity provider is fine, it's Flexential that looks incredibly opaque and non-communicative in a stressful situation. While Cloudflare should have been better prepared for this, it seems to be amateur hour in that particular Portland data-center. Other customers (Dreamhost, etc) were impacted too, and I can't imagine they don't also have some very pointed questions.

Sure, but DreamHost recovered fully within 12 hours [1], Cloudflare took almost 2 days [2]

[1] https://www.dreamhoststatus.com/pages/incident/575f0f6068263... [2] https://www.cloudflarestatus.com/incidents/hm7491k53ppg

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#110
post #93

Earlier quoted context omitted.

Wikipedia's entry for Hillsboro: "Elevation 194 ft (60 m)" Between that, and being ~50 miles inland - I'd say there's ~zero threat of Cascadia quakes or tsunamis directly knocking out those DC's. (Yeah, larger-scale infrastructure and social order could still be killers.) OTOH - Mt. St. Helens is about 60 miles NNE of Hillsboro. If that really went boom, and the wind was right...how many cm's of dry volcanic ash can…

I was never worried about the tsunami. Okay, maybe not gone, but I wouldn't say it would be operational. https://www.oregon.gov/oem/Documents/Cascadia_Rising_Exercis... 50% of roads and near 75% of bridges damaged on the west coast and the I5 corridor. Refer to PDF page #93 where over 70% of power generation is highly damaged on the I5 corridor and 60% in the coastal areas with 0% undamaged. Highly damaged - "Extensi…

Good backups generators at their colo's could handle the lack of utility power for days to weeks. More & better generators could be hauled in and connected.*

The two big problems I'd see would be (1) Social Order and (2) Internet Connectivity. DC's are not fortresses, and internet backbone fibers/routers/etc. are distributed & kinda fragile.

*After all the large-scale power outages & near-outages of recent decades, Cloudflare has no excuse if they lack really-good backup generators at critical facilities. And with their size, Cloudflare must support enough "critical during major disaster" internet services to actually get such generators.

Post reply on HN