Live data from Hacker News

Post Mortem on Cloudflare Control Plane and Analytics Outage

blog.cloudflare.com

21–30 of 241 posts

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#21
> While most of our critical control plane systems had been migrated to the high availability cluster, some services, especially for some newer products, had not yet been added to the high availability cluster.

> The other two data centers running in the area would take over responsibility for the high availability cluster and keep critical services online. Generally that worked as planned. Unfortunately, we discovered that a subset of services that were supposed to be on the high availability cluster had dependencies on services exclusively running in PDX-04.

> A handful of products did not properly get stood up on our disaster recovery sites. These tended to be newer products where we had not fully implemented and tested a disaster recovery procedure.

So the root cause for the outage was that they relied on a single data center. I find that pretty embarrassing for a company like Cloudflare, which powers such relevant parts of the internet.

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#24
post #15

I love how thorough Cloudflare post mortem’s are. Reading the frank, transparent explanations are like a breath of fresh air compared to the obfuscation of nearly every other company comm’s strategy. We were affected but it’s blog posts like these that make me never want to move away. Everyone makes mistakes. Everyone has bad days. It’s how you react afterwards that makes the difference.

I agree, but I also think that for security purposes they should leave out extraneous detail. Also, I know they want to hold their suppliers accountable, but I would hold off pointing fingers. It doesn't really improve behavior, and it makes incentives worse. I really appreciate that they're going to fix the process errors here. But as they suggested, there's a tension between moving fast and being sure. This is typi…

What "security purposes"? Good security isn't based on ignorance of a system, it is on the system being good. We create a self fulfilling prophecy when we hide security practices because what happens is then very few will properly implement their security. Openness is necessary for learning.

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#25
post #21

> While most of our critical control plane systems had been migrated to the high availability cluster, some services, especially for some newer products, had not yet been added to the high availability cluster. > The other two data centers running in the area would take over responsibility for the high availability cluster and keep critical services online. Generally that worked as planned. Unfortunately, we discover…

> I find that pretty embarrassing for a company like Cloudflare, which powers such relevant parts of the internet.

Absolute lack of faith in cloudflare rn.

This is amateur hour stuff.

It's especially egregious that these are new services that were rolled out without HA.

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#26
post #21

> While most of our critical control plane systems had been migrated to the high availability cluster, some services, especially for some newer products, had not yet been added to the high availability cluster. > The other two data centers running in the area would take over responsibility for the high availability cluster and keep critical services online. Generally that worked as planned. Unfortunately, we discover…

[flagged]

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#27
post #18

Earlier quoted context omitted.

Of the incident? Someone on my team called me about 30 minutes after it started. It was challenging for me to stay on top of because it was also the same day as our Q3 earnings call. But team kept me informed throughout the day. I helped where I could. And they handled a very difficult situation very well. That said, lots we can learn from and improve.

What I find bizarre is that the Cloudflare share price jumped when the outage happend! Having read the post mortem, I do not think it could have been handled any better. I think the decision to extend the outage in order to provide rest was absolutely correct. I always enjoy reading these reports from Cloudflare as they are the best in the business.

There's a class of investor (and their trade bots presumably) that sees outrage over a service outage as proof the provider is now mission critical, hence able to "extract value" from the market.

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#29
> However, we had never tested fully taking the entire PDX-04 facility offline.

That is a painful lesson, but unless you are physically powering off the dc or at least disconnecting the network from the outside world you are not testing a real disaster.

You can point fingers at the facility operators, but at the end of the day you have to be able to recover from a dc going completely offline and maybe never coming back. Mother Nature may wipe it off the face of the earth.

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#30
post #21

> While most of our critical control plane systems had been migrated to the high availability cluster, some services, especially for some newer products, had not yet been added to the high availability cluster. > The other two data centers running in the area would take over responsibility for the high availability cluster and keep critical services online. Generally that worked as planned. Unfortunately, we discover…

> I find that pretty embarrassing for a company like Cloudflare, which powers such relevant parts of the internet. Absolute lack of faith in cloudflare rn. This is amateur hour stuff. It's especially egregious that these are new services that were rolled out without HA.

?

Tbh. As far as I can see, their data plane worked at the edge.

Cloudflare released a lot of new products and the ones that affected were: streams, new image upload and logpush.

Their control plane was bad though. But since most products worked, that's more redundancy than most products.

The proposed solution is simple:

- GA requires to be in the high availability cluster

- test entire DC outages

Post reply on HN