Live data from Hacker News

Post Mortem on Cloudflare Control Plane and Analytics Outage

blog.cloudflare.com

131–140 of 241 posts

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#131
post #21

> While most of our critical control plane systems had been migrated to the high availability cluster, some services, especially for some newer products, had not yet been added to the high availability cluster. > The other two data centers running in the area would take over responsibility for the high availability cluster and keep critical services online. Generally that worked as planned. Unfortunately, we discover…

> While most of our critical control plane systems had been migrated to the high availability cluster, some services, especially for some newer products, had not yet been added to the high availability cluster.

It's amazing that they don't have standards that mandate all new systems to use HA from the beginning.

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#132

I love how thorough Cloudflare post mortem’s are. Reading the frank, transparent explanations are like a breath of fresh air compared to the obfuscation of nearly every other company comm’s strategy. We were affected but it’s blog posts like these that make me never want to move away. Everyone makes mistakes. Everyone has bad days. It’s how you react afterwards that makes the difference.

> Everyone makes mistakes. Everyone has bad days.

The issue is when you start having bad days every other day though. We use and depend on CloudFlare Images heavily, it has now been down more than 67 hours over the last 30 days (22h on October 9th, 42h Nov 2 - Nov 4 and a sprinkle of ~hour long outages in between). That's 90.6% availability over the last month.

Transparency is a great differentiator between providers that are fighting in the 99.9% availability range, but when you are hanging on for dear life to stay above the one 9 availability, it doesn't matter.

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#133
post #21

> While most of our critical control plane systems had been migrated to the high availability cluster, some services, especially for some newer products, had not yet been added to the high availability cluster. > The other two data centers running in the area would take over responsibility for the high availability cluster and keep critical services online. Generally that worked as planned. Unfortunately, we discover…

[flagged]

Sounds like chatgpt doesn't want your business and tuned thier cloudflare settings accordingly. Conveniently cloudflare is getting the blame, which is presumably part of what they're paying for.

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#134
Its weird that upon reading this post, I have less confidence in Cloudflare. They basically browbeat Flexential for behaving unprofessional, which, yes, they probably did. However the fact that this causes entire systems that people rely on to go down is a massive redundancy failure on Cloudflares part, you should be able to nuke one of these datacentres and still maintain services.

Very worrying is they start by stating their intended design:

> Cloudflare's control plane and analytics systems run primarily on servers in three data centers around Hillsboro, Oregon

You need way more geographic dispersion than that, this control pane is used by people across the world. We are still on the intended design, not the flawed implementation by the way, which is wild to me.

> This is a system design that we began implementing four years ago. While most of our critical control plane systems had been migrated to the high availability cluster, some services, especially for some newer products, had not yet been added to the high availability cluster.

I don't understand why this would ever be done in this way. If Cloudflare is making a new product for consumers shouldn't redundant design be at the forefront here? I am surprised that it was even an option. For the record I do use Cloudflare for certain systems and I use it because I assume it has great failovers if events like this occur making me not have to worry about these eventualities, but now I will be reconsidering this, how do I actually know my cloudflare workers are safe from these design decisions?

> When services were turned up there, we experienced a thundering herd problem where the API calls that had been failing overwhelmed our services.

Yeh I'll bet, its because Cloudflares core design is not redundant.

Really disappointed in this blog post trying to shift the blame to Flexential when this slapdash architecture should be the main problem on show. As a customer I don't care if Flexential disappears in an earthquake tomorrow, I expect Cloudflare to handle it gracefully.

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#135

Interesting choice to spend the bulk of the article publicly shifting blame to a vendor by name and speculating on their root cause. Also an interesting choice to publicly call out that you're a whale in the facility and include an electrical diagram clearly marked Confidential by your vendor in the postmortem. Honestly, this is rather unprofessional. I understand and support explaining what triggered the event and g…

If Flexential and PGE aren't sharing information or otherwise cooperating as much as Cloudflare might like, then going public with some speculation might be an attempt at applying some pressure to get to the bottom of what happened.

It might also be an effort to get out in front of the story before someone else does the speculating.

In any case, with at least three parties involved, with multiple interconnected systems… if Cloudflare is going to effectively anticipate this cluster of failure modes in future design decisions, it's reasonable for them to want to know what happened all the way down.

Edit to add: I for one am grateful for the information Cloudflare is sharing.

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#136
post #81

Earlier quoted context omitted.

Well, they did distribute their systems. Some were in the running DC, some were not ;)

Their uptime was eventually consistent

haha. The control plane was eventually consistent after 3 days

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#137
post #40

I am really upset about this situation on behalf of CF however why don't they think about generating their own electricity with renewable energy sources?

> why don't they think about generating their own electricity with renewable energy sources?

How exactly do you imagine that working while inside a data center operated by a third party?

It's not like they let you stick some solar panels on the roof and run an extension cord to your rack.

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#138

Its weird that upon reading this post, I have less confidence in Cloudflare. They basically browbeat Flexential for behaving unprofessional, which, yes, they probably did. However the fact that this causes entire systems that people rely on to go down is a massive redundancy failure on Cloudflares part, you should be able to nuke one of these datacentres and still maintain services. Very worrying is they start by sta…

Is the Hillsboro thing is about latency?

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#139
appreciate the status updates and the quick report. as a person that handles the tiniest datacenter, it's impossible to predict every potential event. best you can do is to recover as quickly as possible and learn the lesson.

believing that this doesn't or can't happen to another vendor is being naive.

it has happened to all of them and it'll happen again. can only hope it's super rare.

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#140
post #54

As someone who was slightly affected by this outage, I personally also find this post-mortem to be lacking. 75% of the post-mortem talks about the power outage at PDX-04 and blames Flexential. Okay, fair - it was a bit of a disaster what was happening there judging from the text. But by end of November 2 (UTC), power was fully restored. It still took ~30 hours according to the post-mortem for Cloudflare to fully reco…

I'm not that surprised at the relative lack of detail, given how quickly they released this; I'm surprised they published this much info so quickly. Calling it a postmortem is a bit of a misnomer, though. I'd expect a full postmortem to have the kind of detail you mention.
Post reply on HN