Live data from Hacker News

Post Mortem on Cloudflare Control Plane and Analytics Outage

blog.cloudflare.com

121–130 of 241 posts

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#121
post #50
post #21

> While most of our critical control plane systems had been migrated to the high availability cluster, some services, especially for some newer products, had not yet been added to the high availability cluster. > The other two data centers running in the area would take over responsibility for the high availability cluster and keep critical services online. Generally that worked as planned. Unfortunately, we discover…

And that this was unironically written in the same post mortem: “We are good at distributed systems.” There’s a lack of awareness there.

Good != infallible

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#122
post #10

Contrary to others here, I find the postmortem a bit lacking. The TLDR is that CF runs in multiple data centers, one went down, and the services that depend on it went down with it. The interesting question would be why those services did depend on a single data center. They are pretty vague about it Cloudflare allows multiple teams to innovate quickly. As such, products often take different paths toward their initia…

>I would look into the specific decisions of the engineers and why they decided to make services depend on just one data center

And the product team defining requirements

And IT/governance/architecture teams for not properly cataloging dependencies

And the sales and marketing team not clearly articulating what they're selling (a beta/early access product that's not HA)

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#124
post #17

> It is not unusual for utilities to ask data centers to drop off the grid when power demands are high and run exclusively on generators. Are the data centers compensated or anything for this? I'd imagine generator-only might cost more in terms of fuel and wear-and-tear/maintinaince/inspections. edit: > DSG allows the local utility to run a data center's generators to help supply additional power to the grid. In exch…

I'm not very well versed in this space but I've been told Progressive Insurance in Cleveland, OH has a similar (sounding) agreement. According to PGE's website, they basically pay for everything https://portlandgeneral.com/save-money/save-money-business/d...

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#125

Interesting choice to spend the bulk of the article publicly shifting blame to a vendor by name and speculating on their root cause. Also an interesting choice to publicly call out that you're a whale in the facility and include an electrical diagram clearly marked Confidential by your vendor in the postmortem. Honestly, this is rather unprofessional. I understand and support explaining what triggered the event and g…

Yeah I agree. The data center should be able to blow up without causing any problems. That's what Cloudflare sells and I'm surprised a data center failure can cause such problems.

Going into such depths on the 3rd party just shows how embarrassing this is for them.

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#126

Earlier quoted context omitted.

> I find that pretty embarrassing for a company like Cloudflare, which powers such relevant parts of the internet. Absolute lack of faith in cloudflare rn. This is amateur hour stuff. It's especially egregious that these are new services that were rolled out without HA.

? Tbh. As far as I can see, their data plane worked at the edge. Cloudflare released a lot of new products and the ones that affected were: streams, new image upload and logpush. Their control plane was bad though. But since most products worked, that's more redundancy than most products. The proposed solution is simple: - GA requires to be in the high availability cluster - test entire DC outages

The combination of "newer products" and then having "our Stream service" as the only named service in the post-mortem is very odd, since Stream is hardly a "newer product". It was launched in 2017 and went GA in 2018[2]. If after 5 years it still didn't have a disaster recovery procedure I find it hard to believe they even considered it.

[1]: https://blog.cloudflare.com/introducing-cloudflare-stream/ [2]: https://www.cloudflare.com/press-releases/2018/cloudflare-st...

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#127
post #21

> While most of our critical control plane systems had been migrated to the high availability cluster, some services, especially for some newer products, had not yet been added to the high availability cluster. > The other two data centers running in the area would take over responsibility for the high availability cluster and keep critical services online. Generally that worked as planned. Unfortunately, we discover…

> I am sorry and embarrassed for this incident and the pain that it caused our customers and our team.

So do they.

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#128

Earlier quoted context omitted.

Maybe, but I think that their "Informed Speculation" section was probably unnecessary. They may or may not be correct, but give Flexential an opportunity to share what actually happened rather than openly guessing on what might have happened. Instead, state the facts you know and move onto your response and lessons learned.

I prefer to have their informed speculation here. Has Flexential provided a similarly detailed, public root cause analysis? If so, maybe we can refer to it. If not, how do you expect us to read it?

It’s only been a couple of business days, and it’s likely that they themselves will need root cause from equipment vendors (and perhaps information from the utility) to fully explain what happened. Perhaps they won’t publish anything, but at least give them an opportunity before trying to do it for them.

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#129

Somewhat amazed at the structure of this article: after first discussing the third-party for 75% of blog post, the first-party recovery efforts were detailed in considerably lesser paragraphs. It’s promising to see a path forward mentioned but I can’t help but wonder why this was published instead of currently acknowledging their failure/circumstances and later on publishing a complete post-mortem after the dust full…

To make sure their stonk doesn’t drop at market open next week. Investors will read this (or get the sound bites) and shrug it off as some vendor issue rather than deep issue that will require months of rework (millions of dollars and thus impacting earnings)

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#130

Well, this is definitely NOT a blameless post-mortem! Obviously I'm joking because they are blaming an external company (Flexential) to which they are surely paying big money for the DC space.

I wonder if CF execs aiming to use this to get out of their long term contract with them?
Post reply on HN