Why the very first step was not to fail over Europe?
My question too, although possibly it seemed as a greater risk first to fail over. BTW, is there any unexpected GDPR implication of that? Assuming that fail over means restoring US backups in EU.
Post Mortem on Cloudflare Control Plane and Analytics Outage
71–80 of 241 posts
Re: Post Mortem on Cloudflare Control Plane and Analytics Outage
#72As someone who was slightly affected by this outage, I personally also find this post-mortem to be lacking. 75% of the post-mortem talks about the power outage at PDX-04 and blames Flexential. Okay, fair - it was a bit of a disaster what was happening there judging from the text. But by end of November 2 (UTC), power was fully restored. It still took ~30 hours according to the post-mortem for Cloudflare to fully reco…
> Also one aspect I am missing here is the lack of communication - especially to Enterprise customers. They blame Flexential for lack of communication, but were the first one not saying anything.
Re: Post Mortem on Cloudflare Control Plane and Analytics Outage
#73At what time were you notified Matt?
Of the incident? Someone on my team called me about 30 minutes after it started. It was challenging for me to stay on top of because it was also the same day as our Q3 earnings call. But team kept me informed throughout the day. I helped where I could. And they handled a very difficult situation very well. That said, lots we can learn from and improve.
Re: Post Mortem on Cloudflare Control Plane and Analytics Outage
#74Honestly, this is rather unprofessional. I understand and support explaining what triggered the event and giving a bit of context, but the focus on your postmortem needs to be on your incident, not your vendor's.
Clearly, a lot went wrong and Flexential needs to do their own postmortem, but Cloudflare doesn't need to make guesses and do it for them, much less publicly.
Re: Post Mortem on Cloudflare Control Plane and Analytics Outage
#75Re: Post Mortem on Cloudflare Control Plane and Analytics Outage
#76Earlier quoted context omitted.
I agree, but I also think that for security purposes they should leave out extraneous detail. Also, I know they want to hold their suppliers accountable, but I would hold off pointing fingers. It doesn't really improve behavior, and it makes incentives worse. I really appreciate that they're going to fix the process errors here. But as they suggested, there's a tension between moving fast and being sure. This is typi…
> I also think that for security purposes they should leave out extraneous detail Disagree completely, it's the frank detail that makes me trust their story.
Re: Post Mortem on Cloudflare Control Plane and Analytics Outage
#77Re: Post Mortem on Cloudflare Control Plane and Analytics Outage
#78Earlier quoted context omitted.
? Tbh. As far as I can see, their data plane worked at the edge. Cloudflare released a lot of new products and the ones that affected were: streams, new image upload and logpush. Their control plane was bad though. But since most products worked, that's more redundancy than most products. The proposed solution is simple: - GA requires to be in the high availability cluster - test entire DC outages
From what I was reading on the status page & customers here on HN, WARP + Zero Trust were also majorly affected, which would be quite impactful for a company using these products for their internal authentication. It's not just streams, image upload & Logpush.