Wow they REALLY buried this important part didn't they! This took a ton of scrolling: "Unfortunately, we discovered that a subset of services that were supposed to be on the high availability cluster had dependencies on services exclusively running in PDX-04." Bingo, there we have it.
In my experience, power is the most common data center failure there is. Often it's the redundant systems that cause the failure.
Nobody cares why a data centre died. It's like complaining one of your nodes in a kubernetes cluster has died, or one of your disks in a raid.
The problem here, which is 100% Cloudflare, is that their systems were not resilient across geography.
That's not true. This is behaviour that would be enough for me to pull the plug working with this DC as this is more than unacceptable.
> if you want to have a good working relationship with What are you disagreeing with OP ? He is talking about how to behave if you continue the relationship not whether to continue it .
The post you're replying to is pointing out that multiple days without reporting out a preliminary root cause analysis is so absurdly below the expected level of service here that it would prompt them to reconsider using the service at all.
2 days is outrageous here, I have to imagine whoever thinks that is acceptable is approaching this from the perspective of a company whose downtime doesn't affect profits.
> I find that pretty embarrassing for a company like Cloudflare, which powers such relevant parts of the internet. Bah, who cares about such unimportant details, what's important is that ~dev velocity~ was reaaally high right until that moment! > We were also far too lax about requiring new products and their associated databases to integrate with the high availability cluster. Cloudflare allows multiple teams to inn…
> Complete and utter management failure. And customers apparently are sold what Cloudflare internally considers to be alpha quality software? This has been my experience with AWS and GCP as well. Assume anything that's under 3 years old is not really GA quality no matter what they say publicly.
> This has been my experience with AWS and GCP as well. Assume anything that's under 3 years old is not really GA quality no matter what they say publicly.
GCP run multi-year betas of services and features, so I'm doubtful there were still things not ironed out for GA. Do you have some examples?
they are 14 years old at this point. aws has what, four years on them?
AWS was the public release of tooling that amazon had been bulding for almost 20 years at that point. Similar story for GCP. All three of them had decades of institutional knowledge and procedures in place around running big services by the time Cloudflare was founded.
> AWS was the public release of tooling that amazon had been bulding for almost 20 years at that point.
No, even at the onset AWS was an entirely-from-the-ground-up build. The only thing it could even be argued to sit on top of was the extremely crufty VMs and physical loadbalancers from the original Prod at that point, and those things were not doing anybody any favors.
> if you want to have a good working relationship with What are you disagreeing with OP ? He is talking about how to behave if you continue the relationship not whether to continue it .
The post you're replying to is pointing out that multiple days without reporting out a preliminary root cause analysis is so absurdly below the expected level of service here that it would prompt them to reconsider using the service at all. 2 days is outrageous here, I have to imagine whoever thinks that is acceptable is approaching this from the perspective of a company whose downtime doesn't affect profits.
Interesting choice to spend the bulk of the article publicly shifting blame to a vendor by name and speculating on their root cause. Also an interesting choice to publicly call out that you're a whale in the facility and include an electrical diagram clearly marked Confidential by your vendor in the postmortem. Honestly, this is rather unprofessional. I understand and support explaining what triggered the event and g…
You are way off here, this is 100% on Flexential, they have a 100% Power SLA, that means the power will always be available, right? They also clearly hadn't performed any checks on the circuit breakers and this is a NEWER facility for them, they also didn't even have HALF of the 10hours for the batteries to charge the generators, they also DEFINITELY should have fully moved to generators during this maintenance, they clearly couldn't because they were MORE than likely assisting PGE. Cloudflare CEO is right on here, you pay for Data Center services to be full redundant, they have 18MW at this location and from what I can see they have (2) feeds? That I can't find? Do they? If (1) feed goes down the 2N they have should kick in and with generators there should be NO issues.
My question too, although possibly it seemed as a greater risk first to fail over. BTW, is there any unexpected GDPR implication of that? Assuming that fail over means restoring US backups in EU.
iirc the GDPR prohibits storing EU data in non-EU servers, not vice-versa
But it does mean all that data is now required to be handled in a compliant fashion
The guy I genuinely feel sorry for is "an unaccompanied technician who had only been on the job for a week".
Regardless of any and all corporate spin on the issue, a newbie was dumped into an event at the worst possible time.
I really really really hope he gets a decent bit of counselling to make sure that is fully aware that the issue with the data centre had NOTHING to do with him unplugging the coffee maker to plug in his recharger for his iPhone. Absolutely nothing at all.