Live data from Hacker News

Post Mortem on Cloudflare Control Plane and Analytics Outage

blog.cloudflare.com

151–160 of 241 posts

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#151

> We are a relatively large customer of the facility, consuming approximately 10 percent of its total capacity. I'm surprised that CF are renting space in colocation facilities. I would have expected a business of their size to have their own DCs. Is this common practice for cloud providers?

I'm a little surprised too; I figured they would have their own DCs for their core control plane servers. Colos for their 300+ PoPs makes sense, though.

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#152
post #149

Wow they REALLY buried this important part didn't they! This took a ton of scrolling: "Unfortunately, we discovered that a subset of services that were supposed to be on the high availability cluster had dependencies on services exclusively running in PDX-04." Bingo, there we have it.

What does PDX-04 mean here? Not familiar with how data centers work.

Read the damn article! It's explained at the top.

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#153
post #149

Wow they REALLY buried this important part didn't they! This took a ton of scrolling: "Unfortunately, we discovered that a subset of services that were supposed to be on the high availability cluster had dependencies on services exclusively running in PDX-04." Bingo, there we have it.

What does PDX-04 mean here? Not familiar with how data centers work.

PDX is the airport code for Portland, Oregon, USA. It's the fourth Portland data center.

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#154
post #54

As someone who was slightly affected by this outage, I personally also find this post-mortem to be lacking. 75% of the post-mortem talks about the power outage at PDX-04 and blames Flexential. Okay, fair - it was a bit of a disaster what was happening there judging from the text. But by end of November 2 (UTC), power was fully restored. It still took ~30 hours according to the post-mortem for Cloudflare to fully reco…

I think they just wanted a quick post-mortem. I'm sure they will add more to the blog later in the year when they implement mitigations.

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#155

Earlier quoted context omitted.

Even "we don't know why our data center is failing, but we're sending a team over to physically investigate now" would have been A+ communication in the moment.

Everything was on the status page since the start? DC related updates: > Update - Power to Cloudflare’s core North America data center has been partially restored. Cloudflare has failed over some core services to a backup data center, which has partially remediated impact. Cloudflare is currently working to restore the remaining affected services and bring the core North America data center back online. Nov 02, 2023…

As an enterprise customer, I would expect a CSM reaching out to us informing us about the impact, getting into more details about any restoration plans and potentially even ETAs or rough prioritization to resolution on them.

In reality, Cloudflare's support team was essentially completely unavailable on Nov 2, leaving only the status page. And for most of the day, the updates on the status page were very sparse except "we are working on it", and "We are still seeing gradual improvements and working to restore full functionality.".

Yet clearer status updates were only giving starting on Nov 3. However, I still don't think I heard anything from support or a CSM during that time.

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#156

Wow they REALLY buried this important part didn't they! This took a ton of scrolling: "Unfortunately, we discovered that a subset of services that were supposed to be on the high availability cluster had dependencies on services exclusively running in PDX-04." Bingo, there we have it.

In my experience, power is the most common data center failure there is. Often it's the redundant systems that cause the failure.

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#157

Wow they REALLY buried this important part didn't they! This took a ton of scrolling: "Unfortunately, we discovered that a subset of services that were supposed to be on the high availability cluster had dependencies on services exclusively running in PDX-04." Bingo, there we have it.

In my experience, power is the most common data center failure there is. Often it's the redundant systems that cause the failure.

And that's completely unrelated to my comment but thanks for the insight

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#158

> We are a relatively large customer of the facility, consuming approximately 10 percent of its total capacity. I'm surprised that CF are renting space in colocation facilities. I would have expected a business of their size to have their own DCs. Is this common practice for cloud providers?

> I'm surprised that CF are renting space in colocation facilities. I would have expected a business of their size to have their own DCs. Is this common practice for cloud providers?

Google for one has both. Some GCP regions [0] are in colos, while others are in places where we already had datacenters [1]. We also use colo facilities for peering (and bandwidth offload + connection termination).

I'm under the impression that most AWS Cloudfront locations are also in colo facilities.

[0] https://cloud.google.com/about/locations

[1] https://www.google.com/about/datacenters/locations/

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#160
While debatably unprofessional to blame your vendor, I found this read to be fascinating. I'm sure there are blog posts that detail how data centers work and fail but it's rare to get that cross over from a software engineering context. It puts into perspective what it takes for an average data center of this class to fail: power outage, generator failure, and then battery loss.
Post reply on HN