Live data from Hacker News

Post Mortem on Cloudflare Control Plane and Analytics Outage

blog.cloudflare.com

141–150 of 241 posts

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#141

Earlier quoted context omitted.

> I also think that for security purposes they should leave out extraneous detail Disagree completely, it's the frank detail that makes me trust their story.

Maybe, but I think that their "Informed Speculation" section was probably unnecessary. They may or may not be correct, but give Flexential an opportunity to share what actually happened rather than openly guessing on what might have happened. Instead, state the facts you know and move onto your response and lessons learned.

Yeah, that part really rubbed me the wrong way. If this was a full postmortem published a couple of weeks after the fact and Flexential still wasn't providing details, I could maybe see including it, but this post is the wrong place and time.

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#142
Wow they REALLY buried this important part didn't they! This took a ton of scrolling:

"Unfortunately, we discovered that a subset of services that were supposed to be on the high availability cluster had dependencies on services exclusively running in PDX-04."

Bingo, there we have it.

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#143

I love how thorough Cloudflare post mortem’s are. Reading the frank, transparent explanations are like a breath of fresh air compared to the obfuscation of nearly every other company comm’s strategy. We were affected but it’s blog posts like these that make me never want to move away. Everyone makes mistakes. Everyone has bad days. It’s how you react afterwards that makes the difference.

I would generally agree with you, but this post mortem was 75% blaming Flexential even though it took them almost two days to recover after power was restored. The power outage should have been a single paragraph and then pivoted - DC failures happen, its part of life. Failing to properly account for and recover from it is where the real learnings for Cloudflare are.

It was more of an incident report. The efforts to get back online were mostly around Flexential, so it makes sense to dive in to their failings. That said, it is clear there were major lapses of judgement around the control plane design since they should be able to withstand an earthquake. That they don't have regular disaster recovery testing of the control plane and its dependencies seems crazy. I wonder if it is more that some of those dependencies they hoped to eliminate and replace with in-house technology and hedged their bets on the risk.

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#144

Wow they REALLY buried this important part didn't they! This took a ton of scrolling: "Unfortunately, we discovered that a subset of services that were supposed to be on the high availability cluster had dependencies on services exclusively running in PDX-04." Bingo, there we have it.

Yep, past the part where they spent a long time blaming the data center and power company.

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#145

Wow they REALLY buried this important part didn't they! This took a ton of scrolling: "Unfortunately, we discovered that a subset of services that were supposed to be on the high availability cluster had dependencies on services exclusively running in PDX-04." Bingo, there we have it.

Nah, if only the data center would've stayed up this wouldn't have been a problem. It's clearly on the data center. /s

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#146
post #15

Earlier quoted context omitted.

I agree, but I also think that for security purposes they should leave out extraneous detail. Also, I know they want to hold their suppliers accountable, but I would hold off pointing fingers. It doesn't really improve behavior, and it makes incentives worse. I really appreciate that they're going to fix the process errors here. But as they suggested, there's a tension between moving fast and being sure. This is typi…

> I also think that for security purposes they should leave out extraneous detail Disagree completely, it's the frank detail that makes me trust their story.

Cloudflare vowed to be extremely transparent since the start of their existence. I'm very happy with the fact they have managed to keep this a core company value under extreme growth. I hope it continues after they reach a stable market cap. It isn't like Google that vowed not to be evil until they got big enough to be susceptible to antitrust regulation and negative incentives related to ad revenue.

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#147

Its weird that upon reading this post, I have less confidence in Cloudflare. They basically browbeat Flexential for behaving unprofessional, which, yes, they probably did. However the fact that this causes entire systems that people rely on to go down is a massive redundancy failure on Cloudflares part, you should be able to nuke one of these datacentres and still maintain services. Very worrying is they start by sta…

Is the Hillsboro thing is about latency?

That may well be part of it, some people were talking about the impact of latency in the outage thread [1].

[1] https://news.ycombinator.com/item?id=38113952

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#148
post #70

Earlier quoted context omitted.

> Those customers were impacted until the DC was back up ( 1-2 hours?) On the config plane. Which was still like ~12+ hours, if we check the status page. >Eg. The status page mentioned the healthchecks not working, while everything was fine with it. There were just no analytics at that time to confirm that. What good is a status page that's lying to you? Especially since CF manually updates it, anyway? >Source: I wat…

? This wasn't about status updates going to discord only. There is literally a discussion section on the discord, named: #general-discussions Not everything was clear in the discord too ( eg. The healthchecks were discussed there), that's not something you want to copy-paste in the status updates... Priority for cloudflare seemed to get everything back up. And what they thought was down, was always mentioned in the s…

Oh, I just looked it up and I thought you mean that CF engineers were giving real time updates there. That's not the case.

However, I still fail to see your argument regarding Zero Trust and not being impacted. The status page literally mentioned that the service was recovered on Nov 3, so I don't understand what you mean by:

>The data plane ( which I mentioned) had no issues.

There's literally a section with "Data plane impact" on all over the status page, and ZT is definitely in the earlier ones. And this is given the fact that status updates on Nov 2 were very sparse until power was restored.

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#149

Wow they REALLY buried this important part didn't they! This took a ton of scrolling: "Unfortunately, we discovered that a subset of services that were supposed to be on the high availability cluster had dependencies on services exclusively running in PDX-04." Bingo, there we have it.

What does PDX-04 mean here? Not familiar with how data centers work.

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#150

I love how thorough Cloudflare post mortem’s are. Reading the frank, transparent explanations are like a breath of fresh air compared to the obfuscation of nearly every other company comm’s strategy. We were affected but it’s blog posts like these that make me never want to move away. Everyone makes mistakes. Everyone has bad days. It’s how you react afterwards that makes the difference.

> Everyone makes mistakes. Everyone has bad days. The issue is when you start having bad days every other day though. We use and depend on CloudFlare Images heavily, it has now been down more than 67 hours over the last 30 days (22h on October 9th, 42h Nov 2 - Nov 4 and a sprinkle of ~hour long outages in between). That's 90.6% availability over the last month. Transparency is a great differentiator between providers…

They are a younger company than these other providers. Microsoft, Google, and AWS had their own growth pains and disasters. Remember when Microsoft deleted all the data (contacts, photos, etc) off all their customers Danger phones by accident and had no backup. Talk about naming their product a self-fulfilling prophecy.
Post reply on HN