Earlier quoted context omitted.
> I also think that for security purposes they should leave out extraneous detail Disagree completely, it's the frank detail that makes me trust their story.
Maybe, but I think that their "Informed Speculation" section was probably unnecessary. They may or may not be correct, but give Flexential an opportunity to share what actually happened rather than openly guessing on what might have happened. Instead, state the facts you know and move onto your response and lessons learned.
Post Mortem on Cloudflare Control Plane and Analytics Outage
141–150 of 241 posts
Re: Post Mortem on Cloudflare Control Plane and Analytics Outage
#142"Unfortunately, we discovered that a subset of services that were supposed to be on the high availability cluster had dependencies on services exclusively running in PDX-04."
Bingo, there we have it.
Re: Post Mortem on Cloudflare Control Plane and Analytics Outage
#143I love how thorough Cloudflare post mortem’s are. Reading the frank, transparent explanations are like a breath of fresh air compared to the obfuscation of nearly every other company comm’s strategy. We were affected but it’s blog posts like these that make me never want to move away. Everyone makes mistakes. Everyone has bad days. It’s how you react afterwards that makes the difference.
I would generally agree with you, but this post mortem was 75% blaming Flexential even though it took them almost two days to recover after power was restored. The power outage should have been a single paragraph and then pivoted - DC failures happen, its part of life. Failing to properly account for and recover from it is where the real learnings for Cloudflare are.
Re: Post Mortem on Cloudflare Control Plane and Analytics Outage
#144Wow they REALLY buried this important part didn't they! This took a ton of scrolling: "Unfortunately, we discovered that a subset of services that were supposed to be on the high availability cluster had dependencies on services exclusively running in PDX-04." Bingo, there we have it.
Re: Post Mortem on Cloudflare Control Plane and Analytics Outage
#145Wow they REALLY buried this important part didn't they! This took a ton of scrolling: "Unfortunately, we discovered that a subset of services that were supposed to be on the high availability cluster had dependencies on services exclusively running in PDX-04." Bingo, there we have it.
Re: Post Mortem on Cloudflare Control Plane and Analytics Outage
#146Earlier quoted context omitted.
I agree, but I also think that for security purposes they should leave out extraneous detail. Also, I know they want to hold their suppliers accountable, but I would hold off pointing fingers. It doesn't really improve behavior, and it makes incentives worse. I really appreciate that they're going to fix the process errors here. But as they suggested, there's a tension between moving fast and being sure. This is typi…
> I also think that for security purposes they should leave out extraneous detail Disagree completely, it's the frank detail that makes me trust their story.
Re: Post Mortem on Cloudflare Control Plane and Analytics Outage
#147Its weird that upon reading this post, I have less confidence in Cloudflare. They basically browbeat Flexential for behaving unprofessional, which, yes, they probably did. However the fact that this causes entire systems that people rely on to go down is a massive redundancy failure on Cloudflares part, you should be able to nuke one of these datacentres and still maintain services. Very worrying is they start by sta…
Is the Hillsboro thing is about latency?
Re: Post Mortem on Cloudflare Control Plane and Analytics Outage
#148Earlier quoted context omitted.
> Those customers were impacted until the DC was back up ( 1-2 hours?) On the config plane. Which was still like ~12+ hours, if we check the status page. >Eg. The status page mentioned the healthchecks not working, while everything was fine with it. There were just no analytics at that time to confirm that. What good is a status page that's lying to you? Especially since CF manually updates it, anyway? >Source: I wat…
? This wasn't about status updates going to discord only. There is literally a discussion section on the discord, named: #general-discussions Not everything was clear in the discord too ( eg. The healthchecks were discussed there), that's not something you want to copy-paste in the status updates... Priority for cloudflare seemed to get everything back up. And what they thought was down, was always mentioned in the s…
However, I still fail to see your argument regarding Zero Trust and not being impacted. The status page literally mentioned that the service was recovered on Nov 3, so I don't understand what you mean by:
>The data plane ( which I mentioned) had no issues.
There's literally a section with "Data plane impact" on all over the status page, and ZT is definitely in the earlier ones. And this is given the fact that status updates on Nov 2 were very sparse until power was restored.
Re: Post Mortem on Cloudflare Control Plane and Analytics Outage
#149Wow they REALLY buried this important part didn't they! This took a ton of scrolling: "Unfortunately, we discovered that a subset of services that were supposed to be on the high availability cluster had dependencies on services exclusively running in PDX-04." Bingo, there we have it.
Re: Post Mortem on Cloudflare Control Plane and Analytics Outage
#150I love how thorough Cloudflare post mortem’s are. Reading the frank, transparent explanations are like a breath of fresh air compared to the obfuscation of nearly every other company comm’s strategy. We were affected but it’s blog posts like these that make me never want to move away. Everyone makes mistakes. Everyone has bad days. It’s how you react afterwards that makes the difference.
> Everyone makes mistakes. Everyone has bad days. The issue is when you start having bad days every other day though. We use and depend on CloudFlare Images heavily, it has now been down more than 67 hours over the last 30 days (22h on October 9th, 42h Nov 2 - Nov 4 and a sprinkle of ~hour long outages in between). That's 90.6% availability over the last month. Transparency is a great differentiator between providers…