Post Mortem on Cloudflare Control Plane and Analytics Outage
211–220 of 241 posts
Re: Post Mortem on Cloudflare Control Plane and Analytics Outage
#212Earlier quoted context omitted.
Everything was on the status page since the start? DC related updates: > Update - Power to Cloudflare’s core North America data center has been partially restored. Cloudflare has failed over some core services to a backup data center, which has partially remediated impact. Cloudflare is currently working to restore the remaining affected services and bring the core North America data center back online. Nov 02, 2023…
I've got no knock on the status page. Cloudflare is disappointed in the lack of notification from their data center provider, and Cloudflare customers are disappointed in the lack of notification from their service provider. Instead of defending what was done and calling that good enough, Cloudflare should use this as an opportunity to commit to reevaluating the strategy for customer outreach during major service fai…
You want Cloudflare to update every customer for an issue that they probably aren't affected with ( except when changing things) ?
Who even does that when you've got so many customers?
That's exactly what why the status page is there:
https://www.cloudflarestatus.com/
The DC obviously didn't have any means to update their customers.
Re: Post Mortem on Cloudflare Control Plane and Analytics Outage
#213Earlier quoted context omitted.
? 1) Were you affected on the data plane? Which product? As far as I can tell, while the outage was in the core dc's. The impact was minor. 2) Both examples were exactly from 2 November. Not 3 November. 3) What method of support did you try? I thought that their support was impacted ( email?). The status page explicitly mentioned to get in contact with your account manager for some config changes on some products, if…
>1) Were you affected on the data plane? Which product? No, but we needed to make urgent changes. >2) Both examples were exactly from 2 November. Not 3 November. Both messages contain no clear messages about remediation and co. They also didn't state clearly which products were failed over. I noticed that at this point I could at least login to the dashboard, but most stuff was still severely broken, and I had no ide…
We are a service provider in ( mostly) Europe.
Our policy ( playbook) in case of an issue is updating the status page as quick as possible and customers can subscribe on RSS.
There was one issue in the past where we wanted to inform the clients. But it's not easy, as only some were impacted and we decided against it.
5 minutes later ( it was out of our hands) it was solved...
Our playbook is too update the status page as soon as possible to inform the clients something is up and we are aware.
There shouldn't be too much info on it, since sometimes you just aren't 100% sure about what's exactly going on.
We also decided that we want provide durations on it, since you then create a commitment that's possibly dependent on external factors.
Tbh. I can completely understand the approach from Cloudflare here. With an issue, support is overwhelmed. That's why you use the status page ASAP.
Technical details happen in the post-mortem. When we can be sure if any data is lost ( normally, there is nothing lost though, but it's possible we need to requeue some actions)
=> this is when we can contact our clients and brought up to date.
Depending on the SLA it's included or eg. Is paid extra ( in a lot of times, an external provider fails and we can fix something from our end, eg. Resending some data)
Re: Post Mortem on Cloudflare Control Plane and Analytics Outage
#214Re: Post Mortem on Cloudflare Control Plane and Analytics Outage
#215Interesting choice to spend the bulk of the article publicly shifting blame to a vendor by name and speculating on their root cause. Also an interesting choice to publicly call out that you're a whale in the facility and include an electrical diagram clearly marked Confidential by your vendor in the postmortem. Honestly, this is rather unprofessional. I understand and support explaining what triggered the event and g…
I actually disagree, and think that the post mortem clearly defines that there were things that were disappointing that happened with the vendor, _as well as_ things that were disappointing that happened internally. I don't think that it's unfair to point out everything in an event that happened; I do think it would be unfair to ignore all the compounding issues that were in the power of the vendor, and just swallow…
That said, the existence of the 480V labeled intermediary does suggest they have a 277/480 V outside system, and a 120/208 V rack-side system.
Re: Post Mortem on Cloudflare Control Plane and Analytics Outage
#216Interesting choice to spend the bulk of the article publicly shifting blame to a vendor by name and speculating on their root cause. Also an interesting choice to publicly call out that you're a whale in the facility and include an electrical diagram clearly marked Confidential by your vendor in the postmortem. Honestly, this is rather unprofessional. I understand and support explaining what triggered the event and g…
(none of this takes away from the mistakes that were wholly theirs that shouldn't have happened and that they should fix)
Re: Post Mortem on Cloudflare Control Plane and Analytics Outage
#217Earlier quoted context omitted.
They are a younger company than these other providers. Microsoft, Google, and AWS had their own growth pains and disasters. Remember when Microsoft deleted all the data (contacts, photos, etc) off all their customers Danger phones by accident and had no backup. Talk about naming their product a self-fulfilling prophecy.
they are 14 years old at this point. aws has what, four years on them?
Similar story for GCP.
All three of them had decades of institutional knowledge and procedures in place around running big services by the time Cloudflare was founded.
Re: Post Mortem on Cloudflare Control Plane and Analytics Outage
#218Re: Post Mortem on Cloudflare Control Plane and Analytics Outage
#219Earlier quoted context omitted.
they are 14 years old at this point. aws has what, four years on them?
AWS was the public release of tooling that amazon had been bulding for almost 20 years at that point. Similar story for GCP. All three of them had decades of institutional knowledge and procedures in place around running big services by the time Cloudflare was founded.
Re: Post Mortem on Cloudflare Control Plane and Analytics Outage
#220Earlier quoted context omitted.
Dunno about that, I've read similar internal postmortems at the FAANG I worked at.
Everywhere I've worked requires a DR drill per service, but I've never seen anything where the whole company shuts down a DC at once across all services. But probably we should. It's an immensely larger coordination problem, but frankly, it's probably the more common failure mode.