Live data from Hacker News

Post Mortem on Cloudflare Control Plane and Analytics Outage

blog.cloudflare.com

171–180 of 241 posts

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#171

Earlier quoted context omitted.

>If Flexential and PGE aren't sharing information or otherwise cooperating as much as Cloudflare might like, then going public with some speculation might be an attempt at applying some pressure to get to the bottom of what happened. It's been 2 days. I doubt PGE or Flexential even have root caused it yet, and even if they have, good communication takes time. You don't throw someone under the bus and smear their name…

> You don't throw someone under the bus and smear their name publicly just because they haven't replied for two days, and you certainly don't start speculating on their behalf. That's bad partnership. 1. When you’re paying them the kind of money I imagine they’re paying and they don’t reply for 2 days, yea that’s crazy if true. I’d expect a client of this size could take to an executive on their personal number. 2. T…

We have no idea what their contract is. But two business days without a reply isn’t exactly a long time. Especially if they are conducting their own investigation and reproduction steps.

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#172
post #106

Earlier quoted context omitted.

Especially since it shouldn't matter why the DC failed — Cloudflare's entire business model is selling services allegedly designed to survive that. 99% of the fault lies with Cloudflare for not being able to do their core job.

In all fairness the rest of the article is about that

Slightly less than half, and the bottom half, so that people just skimming over it will mostly remember the DC operators' problems, not Cloudflare's own. This is very deliberately manipulative.

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#173
For me it's basically summed up as "we didn't test turning the power off" and making sure things worked the way we planned.

Yes it is hard and very expensive to do these types of tests. And doing it regularly is even more $$$ and time.

As most customers we seem to be okay with a cheap price hidden behind a facade of "high availability" since I don't really want to pay for true HA. Because if I knew the real cost it would be too expensive.

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#174
post #148

Earlier quoted context omitted.

? This wasn't about status updates going to discord only. There is literally a discussion section on the discord, named: #general-discussions Not everything was clear in the discord too ( eg. The healthchecks were discussed there), that's not something you want to copy-paste in the status updates... Priority for cloudflare seemed to get everything back up. And what they thought was down, was always mentioned in the s…

Oh, I just looked it up and I thought you mean that CF engineers were giving real time updates there. That's not the case. However, I still fail to see your argument regarding Zero Trust and not being impacted. The status page literally mentioned that the service was recovered on Nov 3, so I don't understand what you mean by: >The data plane ( which I mentioned) had no issues. There's literally a section with "Data p…

We don't use zero trust atm. So, I can't know for sure.

What I mentioned, was what I've seen passing by in the channel at the time.

I also saw no incoming help requests for zero trust tbh ( did some community help)

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#175

Earlier quoted context omitted.

>If Flexential and PGE aren't sharing information or otherwise cooperating as much as Cloudflare might like, then going public with some speculation might be an attempt at applying some pressure to get to the bottom of what happened. It's been 2 days. I doubt PGE or Flexential even have root caused it yet, and even if they have, good communication takes time. You don't throw someone under the bus and smear their name…

> You don't throw someone under the bus and smear their name publicly just because they haven't replied for two days, and you certainly don't start speculating on their behalf. That's bad partnership. 1. When you’re paying them the kind of money I imagine they’re paying and they don’t reply for 2 days, yea that’s crazy if true. I’d expect a client of this size could take to an executive on their personal number. 2. T…

They aren't telling the facts as they know them. Cloudflare themselves say that the information in the article is "speculation" (the article literally uses that term).

Publicly casting blame based on speculation isn't something you do to someone that you want to have a good working relationship with, no matter how much money you pay them.

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#176
post #106

Earlier quoted context omitted.

Especially since it shouldn't matter why the DC failed — Cloudflare's entire business model is selling services allegedly designed to survive that. 99% of the fault lies with Cloudflare for not being able to do their core job.

In all fairness the rest of the article is about that

So why spend so much time trying to shift blame to the vendor? They could've just started the article with something like:

> Due to circumstances beyond our control the DC lost all power. We are still working with our vendors to investigate the cause. While such a failure should not have been possible, our systems are supposed to tolerate a complete loss of a DC.

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#177
post #150

Earlier quoted context omitted.

> Everyone makes mistakes. Everyone has bad days. The issue is when you start having bad days every other day though. We use and depend on CloudFlare Images heavily, it has now been down more than 67 hours over the last 30 days (22h on October 9th, 42h Nov 2 - Nov 4 and a sprinkle of ~hour long outages in between). That's 90.6% availability over the last month. Transparency is a great differentiator between providers…

They are a younger company than these other providers. Microsoft, Google, and AWS had their own growth pains and disasters. Remember when Microsoft deleted all the data (contacts, photos, etc) off all their customers Danger phones by accident and had no backup. Talk about naming their product a self-fulfilling prophecy.

they are 14 years old at this point. aws has what, four years on them?

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#178
post #150

Earlier quoted context omitted.

> Everyone makes mistakes. Everyone has bad days. The issue is when you start having bad days every other day though. We use and depend on CloudFlare Images heavily, it has now been down more than 67 hours over the last 30 days (22h on October 9th, 42h Nov 2 - Nov 4 and a sprinkle of ~hour long outages in between). That's 90.6% availability over the last month. Transparency is a great differentiator between providers…

They are a younger company than these other providers. Microsoft, Google, and AWS had their own growth pains and disasters. Remember when Microsoft deleted all the data (contacts, photos, etc) off all their customers Danger phones by accident and had no backup. Talk about naming their product a self-fulfilling prophecy.

Cloudflare is fourteen years old

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#179

Its weird that upon reading this post, I have less confidence in Cloudflare. They basically browbeat Flexential for behaving unprofessional, which, yes, they probably did. However the fact that this causes entire systems that people rely on to go down is a massive redundancy failure on Cloudflares part, you should be able to nuke one of these datacentres and still maintain services. Very worrying is they start by sta…

I'm also a bit surprised about Hillsboro. The FEMA is assuming that when (not if) The Big One hits, everything west of I-5 is going to be toast.

Is placing the entirety of such a critical cluster in a known earthquake and tsunami zone a good idea? It looks like their disaster recovery to Europe didn't really work either...

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#180
post #48

Earlier quoted context omitted.

> I find that pretty embarrassing for a company like Cloudflare, which powers such relevant parts of the internet. Bah, who cares about such unimportant details, what's important is that ~dev velocity~ was reaaally high right until that moment! > We were also far too lax about requiring new products and their associated databases to integrate with the high availability cluster. Cloudflare allows multiple teams to inn…

> Complete and utter management failure Too strong. A failure certainly, but painting this as the worst possible management failure is kind of silly.

To be honest if you take the circumstances and them spending half of their post-mortem blaming the vendor, it does look like a total shitshow.
Post reply on HN