> While most of our critical control plane systems had been migrated to the high availability cluster, some services, especially for some newer products, had not yet been added to the high availability cluster. > The other two data centers running in the area would take over responsibility for the high availability cluster and keep critical services online. Generally that worked as planned. Unfortunately, we discover…
And that this was unironically written in the same post mortem: “We are good at distributed systems.” There’s a lack of awareness there.
Post Mortem on Cloudflare Control Plane and Analytics Outage
121–130 of 241 posts
Re: Post Mortem on Cloudflare Control Plane and Analytics Outage
#122Contrary to others here, I find the postmortem a bit lacking. The TLDR is that CF runs in multiple data centers, one went down, and the services that depend on it went down with it. The interesting question would be why those services did depend on a single data center. They are pretty vague about it Cloudflare allows multiple teams to innovate quickly. As such, products often take different paths toward their initia…
And the product team defining requirements
And IT/governance/architecture teams for not properly cataloging dependencies
And the sales and marketing team not clearly articulating what they're selling (a beta/early access product that's not HA)
Re: Post Mortem on Cloudflare Control Plane and Analytics Outage
#123True, not snippy: I found interesting that their automated billing emails seemed to arrive right on time.
Re: Post Mortem on Cloudflare Control Plane and Analytics Outage
#124> It is not unusual for utilities to ask data centers to drop off the grid when power demands are high and run exclusively on generators. Are the data centers compensated or anything for this? I'd imagine generator-only might cost more in terms of fuel and wear-and-tear/maintinaince/inspections. edit: > DSG allows the local utility to run a data center's generators to help supply additional power to the grid. In exch…
Re: Post Mortem on Cloudflare Control Plane and Analytics Outage
#125Interesting choice to spend the bulk of the article publicly shifting blame to a vendor by name and speculating on their root cause. Also an interesting choice to publicly call out that you're a whale in the facility and include an electrical diagram clearly marked Confidential by your vendor in the postmortem. Honestly, this is rather unprofessional. I understand and support explaining what triggered the event and g…
Going into such depths on the 3rd party just shows how embarrassing this is for them.
Re: Post Mortem on Cloudflare Control Plane and Analytics Outage
#126Earlier quoted context omitted.
> I find that pretty embarrassing for a company like Cloudflare, which powers such relevant parts of the internet. Absolute lack of faith in cloudflare rn. This is amateur hour stuff. It's especially egregious that these are new services that were rolled out without HA.
? Tbh. As far as I can see, their data plane worked at the edge. Cloudflare released a lot of new products and the ones that affected were: streams, new image upload and logpush. Their control plane was bad though. But since most products worked, that's more redundancy than most products. The proposed solution is simple: - GA requires to be in the high availability cluster - test entire DC outages
[1]: https://blog.cloudflare.com/introducing-cloudflare-stream/ [2]: https://www.cloudflare.com/press-releases/2018/cloudflare-st...
Re: Post Mortem on Cloudflare Control Plane and Analytics Outage
#127> While most of our critical control plane systems had been migrated to the high availability cluster, some services, especially for some newer products, had not yet been added to the high availability cluster. > The other two data centers running in the area would take over responsibility for the high availability cluster and keep critical services online. Generally that worked as planned. Unfortunately, we discover…
So do they.
Re: Post Mortem on Cloudflare Control Plane and Analytics Outage
#128Earlier quoted context omitted.
Maybe, but I think that their "Informed Speculation" section was probably unnecessary. They may or may not be correct, but give Flexential an opportunity to share what actually happened rather than openly guessing on what might have happened. Instead, state the facts you know and move onto your response and lessons learned.
I prefer to have their informed speculation here. Has Flexential provided a similarly detailed, public root cause analysis? If so, maybe we can refer to it. If not, how do you expect us to read it?
Re: Post Mortem on Cloudflare Control Plane and Analytics Outage
#129Somewhat amazed at the structure of this article: after first discussing the third-party for 75% of blog post, the first-party recovery efforts were detailed in considerably lesser paragraphs. It’s promising to see a path forward mentioned but I can’t help but wonder why this was published instead of currently acknowledging their failure/circumstances and later on publishing a complete post-mortem after the dust full…
Re: Post Mortem on Cloudflare Control Plane and Analytics Outage
#130Well, this is definitely NOT a blameless post-mortem! Obviously I'm joking because they are blaming an external company (Flexential) to which they are surely paying big money for the DC space.