Post Mortem on Cloudflare Control Plane and Analytics Outage
191–200 of 241 posts
Re: Post Mortem on Cloudflare Control Plane and Analytics Outage
#192Re: Post Mortem on Cloudflare Control Plane and Analytics Outage
#193Earlier quoted context omitted.
Your operational reviews must be lacking at AWS then (surprise surprise) then because there are so many instances where something will be released in alpha yet the documentation will still be outdated, stale and incorrect LOL.
I think you misunderstand what's being talked about in this thread. "Operations" in this context has nothing to do with external-facing documentation, and instead refers to the resilience of the service and ensuring it doesn't for example, stop working when a single data center experiences a power outage.
Re: Post Mortem on Cloudflare Control Plane and Analytics Outage
#194Re: Post Mortem on Cloudflare Control Plane and Analytics Outage
#195Earlier quoted context omitted.
They aren't telling the facts as they know them. Cloudflare themselves say that the information in the article is "speculation" (the article literally uses that term). Publicly casting blame based on speculation isn't something you do to someone that you want to have a good working relationship with, no matter how much money you pay them.
If you actually worked with datacenters you'd understand that what PGE and Flexential is unacceptable as well
UPS failing early sounds like it may be a battery maintenance issue.
Re: Post Mortem on Cloudflare Control Plane and Analytics Outage
#196We should absolutely blame them, just as "victims" of ransomware should be blamed. Hardening against system failure is the same process as security hardening.
Re: Post Mortem on Cloudflare Control Plane and Analytics Outage
#197Earlier quoted context omitted.
> You don't throw someone under the bus and smear their name publicly just because they haven't replied for two days, and you certainly don't start speculating on their behalf. That's bad partnership. 1. When you’re paying them the kind of money I imagine they’re paying and they don’t reply for 2 days, yea that’s crazy if true. I’d expect a client of this size could take to an executive on their personal number. 2. T…
They aren't telling the facts as they know them. Cloudflare themselves say that the information in the article is "speculation" (the article literally uses that term). Publicly casting blame based on speculation isn't something you do to someone that you want to have a good working relationship with, no matter how much money you pay them.
Re: Post Mortem on Cloudflare Control Plane and Analytics Outage
#198> Our team was all-hands-on-deck and had worked all day on the emergency, so I made the call that most of us should get some rest and start the move back to PDX-04 in the morning. That decision delayed our full recovery, but I believe made it less likely that we’d compound this situation with additional mistakes. I liked this - the human element is underemphasised often in these kinds of reports, and trying to fix a…
I’m curious, have these plans ever been tested in a real incident? Like Mike Tyson says, everyone has a plan until they get punched in the face.
If you don't do that, you're still going to be scrambling.
Re: Post Mortem on Cloudflare Control Plane and Analytics Outage
#199Earlier quoted context omitted.
As an enterprise customer, I would expect a CSM reaching out to us informing us about the impact, getting into more details about any restoration plans and potentially even ETAs or rough prioritization to resolution on them. In reality, Cloudflare's support team was essentially completely unavailable on Nov 2, leaving only the status page. And for most of the day, the updates on the status page were very sparse excep…
? 1) Were you affected on the data plane? Which product? As far as I can tell, while the outage was in the core dc's. The impact was minor. 2) Both examples were exactly from 2 November. Not 3 November. 3) What method of support did you try? I thought that their support was impacted ( email?). The status page explicitly mentioned to get in contact with your account manager for some config changes on some products, if…
No, but we needed to make urgent changes.
>2) Both examples were exactly from 2 November. Not 3 November.
Both messages contain no clear messages about remediation and co. They also didn't state clearly which products were failed over. I noticed that at this point I could at least login to the dashboard, but most stuff was still severely broken, and I had no idea whether changes with the few semi-functional components were actually applied or not.
Updates to single products with a more clear status were given only at the end of November 2nd (UTC).
(Also one of the message states data centres - not just data center. Not sure what happened there).
>3) What method of support did you try? I thought that their support was impacted ( email?).
Emergency line + contacting our CSM. The emergency line was shut down and replaced with voice mail (WTF?), and our CSM did not reply at all (or the message somehow made it to the wrong person, I'll find out next week, I guess).
So in our case, the communication was essentially non-existent, even though I raised a support case (or wanted to).
>4) I have never heard of Enterprise customers being contacted by a cloud company during an outage. Which company does that? Do you have an example?
I can remember of Datadog reaching out to us for their 2023-03-08 incident. Not sure if it was just our CSM being nice or someone did a support request on another communication channel, but looking back in history that came without asking + the post mortem. Same case when stuff happens such as vulnerabilities in one of their packages, they reach out to us proactively and notify us.
To be fair, this is a bit of a wishlist and definitely not necessary for a 30 minutes hickup, but for a 2 day outage... I don't know.
At the bare minimum, I'd expect at least their support team to be replying and not shutting down the communication channels.
>5) I would think it's absolutely a nogo to contact every preemptively Enterprise customer with: "hey, the product works, but if you change xyz, atm that doesn't.".
I don't know... At least at the time I raise an urgent support case about an issue, I expect to be kept up-to-date.
> Since most customers weren't affected and some others were minorly impacted.
What does it mean they were not affected? Yes, their core service was still functioning (thank god - after all they advertise a 100% (!) SLA on that), but you can see on same Discord channel you mentioned people failing to renew TLS certificates, people couldn't make Vercel deployments and more. So it did affect quite a bunch of downstream customers in their products, and they might also sell SLAs to their customers...
I cannot really comment on whether that just affected us, or if other customers had better support experiences here.
But I expect better in terms of communication here. Doesn't have to be as outreaching as I did in my last message, but stuff like shutting down the emergency line and not giving any comment is not really acceptable for an Enterprise contract.
Re: Post Mortem on Cloudflare Control Plane and Analytics Outage
#200Interesting choice to spend the bulk of the article publicly shifting blame to a vendor by name and speculating on their root cause. Also an interesting choice to publicly call out that you're a whale in the facility and include an electrical diagram clearly marked Confidential by your vendor in the postmortem. Honestly, this is rather unprofessional. I understand and support explaining what triggered the event and g…
And even though Cloudflare did put some of the blame, as it were, on the vendor, the post mortem recognizes that Cloudflare wasn't doing their due diligence on their vendor's maintenance and upkeep to verify that the state of the vendor's equipment is the same as the day they signed on. And that's ignoring a huge focus of the post mortem where they admit guilt at not knowing or not changing the fact that Kafka and Clickhouse were only in that datacenter.
Furthermore, we do not know that Cloudflare didn't get the vendor's blessing to submit that diagram to their post mortem. You're assuming they didn't. But for what it's worth as someone that has worked in datacenters, none of this is all that proprietary. Their business isn't hurt because this came out. This is a fairly standard (and frankly simplified for business folk) diagram of what any decently engineered datacenter building would operate like. There's no magic sauce in here that other datacenter companies are going to steal to put Flexential out of business. If you work for a datacenter company that doesn't already have any of this, you should write a check to Flexential or their electrical engineers for a consultancy.
And finally, the things that Cloudflare speculated on were things like, to paraphrase, "we know that a transformer failed, and we believe that its purpose was to step down the voltage that the utility company was running into the datacenter." Which, if you have basic electrical engineering knowledge, just makes sense. The utility company is delivering 12470 volts, of course that needs to be stepped down, somewhere along the way, probably multiple times, before it ends up coming through the 210 volt rack PDUs. I'm willing to accept that guess in the absence of facts from the vendor while they're still being tight lipped.
However, that's not to say I'm totally satisfied by this post mortem either. I am also interested in hearing what decisions led to them leaving Kafka and Clickhouse in a state of non-redundancy (at least at the datacenter level) or how they could have not known about it. Detail was left out there, for sure.