> We are a relatively large customer of the facility, consuming approximately 10 percent of its total capacity. I'm surprised that CF are renting space in colocation facilities. I would have expected a business of their size to have their own DCs. Is this common practice for cloud providers?
Post Mortem on Cloudflare Control Plane and Analytics Outage
151–160 of 241 posts
Re: Post Mortem on Cloudflare Control Plane and Analytics Outage
#152Wow they REALLY buried this important part didn't they! This took a ton of scrolling: "Unfortunately, we discovered that a subset of services that were supposed to be on the high availability cluster had dependencies on services exclusively running in PDX-04." Bingo, there we have it.
What does PDX-04 mean here? Not familiar with how data centers work.
Re: Post Mortem on Cloudflare Control Plane and Analytics Outage
#153Wow they REALLY buried this important part didn't they! This took a ton of scrolling: "Unfortunately, we discovered that a subset of services that were supposed to be on the high availability cluster had dependencies on services exclusively running in PDX-04." Bingo, there we have it.
What does PDX-04 mean here? Not familiar with how data centers work.
Re: Post Mortem on Cloudflare Control Plane and Analytics Outage
#154As someone who was slightly affected by this outage, I personally also find this post-mortem to be lacking. 75% of the post-mortem talks about the power outage at PDX-04 and blames Flexential. Okay, fair - it was a bit of a disaster what was happening there judging from the text. But by end of November 2 (UTC), power was fully restored. It still took ~30 hours according to the post-mortem for Cloudflare to fully reco…
Re: Post Mortem on Cloudflare Control Plane and Analytics Outage
#155Earlier quoted context omitted.
Even "we don't know why our data center is failing, but we're sending a team over to physically investigate now" would have been A+ communication in the moment.
Everything was on the status page since the start? DC related updates: > Update - Power to Cloudflare’s core North America data center has been partially restored. Cloudflare has failed over some core services to a backup data center, which has partially remediated impact. Cloudflare is currently working to restore the remaining affected services and bring the core North America data center back online. Nov 02, 2023…
In reality, Cloudflare's support team was essentially completely unavailable on Nov 2, leaving only the status page. And for most of the day, the updates on the status page were very sparse except "we are working on it", and "We are still seeing gradual improvements and working to restore full functionality.".
Yet clearer status updates were only giving starting on Nov 3. However, I still don't think I heard anything from support or a CSM during that time.
Re: Post Mortem on Cloudflare Control Plane and Analytics Outage
#156Wow they REALLY buried this important part didn't they! This took a ton of scrolling: "Unfortunately, we discovered that a subset of services that were supposed to be on the high availability cluster had dependencies on services exclusively running in PDX-04." Bingo, there we have it.
Re: Post Mortem on Cloudflare Control Plane and Analytics Outage
#157Wow they REALLY buried this important part didn't they! This took a ton of scrolling: "Unfortunately, we discovered that a subset of services that were supposed to be on the high availability cluster had dependencies on services exclusively running in PDX-04." Bingo, there we have it.
In my experience, power is the most common data center failure there is. Often it's the redundant systems that cause the failure.
Re: Post Mortem on Cloudflare Control Plane and Analytics Outage
#158> We are a relatively large customer of the facility, consuming approximately 10 percent of its total capacity. I'm surprised that CF are renting space in colocation facilities. I would have expected a business of their size to have their own DCs. Is this common practice for cloud providers?
Google for one has both. Some GCP regions [0] are in colos, while others are in places where we already had datacenters [1]. We also use colo facilities for peering (and bandwidth offload + connection termination).
I'm under the impression that most AWS Cloudfront locations are also in colo facilities.