Live data from Hacker News

Post Mortem on Cloudflare Control Plane and Analytics Outage

blog.cloudflare.com

71–80 of 241 posts

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#71
post #44

Why the very first step was not to fail over Europe?

My question too, although possibly it seemed as a greater risk first to fail over. BTW, is there any unexpected GDPR implication of that? Assuming that fail over means restoring US backups in EU.

iirc the GDPR prohibits storing EU data in non-EU servers, not vice-versa

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#72
post #65
post #54

As someone who was slightly affected by this outage, I personally also find this post-mortem to be lacking. 75% of the post-mortem talks about the power outage at PDX-04 and blames Flexential. Okay, fair - it was a bit of a disaster what was happening there judging from the text. But by end of November 2 (UTC), power was fully restored. It still took ~30 hours according to the post-mortem for Cloudflare to fully reco…

> Also one aspect I am missing here is the lack of communication - especially to Enterprise customers. They blame Flexential for lack of communication, but were the first one not saying anything.

Even "we don't know why our data center is failing, but we're sending a team over to physically investigate now" would have been A+ communication in the moment.

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#73

At what time were you notified Matt?

Of the incident? Someone on my team called me about 30 minutes after it started. It was challenging for me to stay on top of because it was also the same day as our Q3 earnings call. But team kept me informed throughout the day. I helped where I could. And they handled a very difficult situation very well. That said, lots we can learn from and improve.

Coincidental timing?

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#74
Interesting choice to spend the bulk of the article publicly shifting blame to a vendor by name and speculating on their root cause. Also an interesting choice to publicly call out that you're a whale in the facility and include an electrical diagram clearly marked Confidential by your vendor in the postmortem.

Honestly, this is rather unprofessional. I understand and support explaining what triggered the event and giving a bit of context, but the focus on your postmortem needs to be on your incident, not your vendor's.

Clearly, a lot went wrong and Flexential needs to do their own postmortem, but Cloudflare doesn't need to make guesses and do it for them, much less publicly.

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#76
post #15

Earlier quoted context omitted.

I agree, but I also think that for security purposes they should leave out extraneous detail. Also, I know they want to hold their suppliers accountable, but I would hold off pointing fingers. It doesn't really improve behavior, and it makes incentives worse. I really appreciate that they're going to fix the process errors here. But as they suggested, there's a tension between moving fast and being sure. This is typi…

> I also think that for security purposes they should leave out extraneous detail Disagree completely, it's the frank detail that makes me trust their story.

Maybe, but I think that their "Informed Speculation" section was probably unnecessary. They may or may not be correct, but give Flexential an opportunity to share what actually happened rather than openly guessing on what might have happened. Instead, state the facts you know and move onto your response and lessons learned.

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#78
post #58

Earlier quoted context omitted.

? Tbh. As far as I can see, their data plane worked at the edge. Cloudflare released a lot of new products and the ones that affected were: streams, new image upload and logpush. Their control plane was bad though. But since most products worked, that's more redundancy than most products. The proposed solution is simple: - GA requires to be in the high availability cluster - test entire DC outages

From what I was reading on the status page & customers here on HN, WARP + Zero Trust were also majorly affected, which would be quite impactful for a company using these products for their internal authentication. It's not just streams, image upload & Logpush.

This was short downtime. But big companies must create own gateway but small just waiting and relying on CF

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#80
Poor doc: You had a high availability 3 data center setup that utterly failed. Why spend the first third of the document blaming your data center operator? The management of the data center facility is outside of your control. You gambled that not appropriately testing your high-availability setup (under your control) would not have consequences. You should absolutely discuss the DC management with your operator, but that's between you and them and doesn't belong in this post mortem.
Post reply on HN