Live data from Hacker News

Post Mortem on Cloudflare Control Plane and Analytics Outage

blog.cloudflare.com

201–210 of 241 posts

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#201
post #155

Earlier quoted context omitted.

As an enterprise customer, I would expect a CSM reaching out to us informing us about the impact, getting into more details about any restoration plans and potentially even ETAs or rough prioritization to resolution on them. In reality, Cloudflare's support team was essentially completely unavailable on Nov 2, leaving only the status page. And for most of the day, the updates on the status page were very sparse excep…

? 1) Were you affected on the data plane? Which product? As far as I can tell, while the outage was in the core dc's. The impact was minor. 2) Both examples were exactly from 2 November. Not 3 November. 3) What method of support did you try? I thought that their support was impacted ( email?). The status page explicitly mentioned to get in contact with your account manager for some config changes on some products, if…

for us as an enterprise customer for many years:

ssl for saas -> custom hostnames are not working for new domains or changes to current ones. also page rules -> redirects are not working for new rules or changes to current rules. which are game-stoppers for our business.

we contacted via enterprise email support + ccing our managers and assigned engineers.

first they try to tell us product is working and sending us some details how to do that,this etc, after a couple of hours later they understand the issue is bigger than they thought and they said "the product is affected by api outage".

then in another email we asked them when this can be solved but only answer we got is "please follow status page for the updates".

and after a day, ssl for saas & ssl services took their places on status page. for a day nobody notices if it's working or not except customers.

so as we understand these emails even the team internally haven't got any idea what is working and what is not!

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#202

Earlier quoted context omitted.

> You don't throw someone under the bus and smear their name publicly just because they haven't replied for two days, and you certainly don't start speculating on their behalf. That's bad partnership. 1. When you’re paying them the kind of money I imagine they’re paying and they don’t reply for 2 days, yea that’s crazy if true. I’d expect a client of this size could take to an executive on their personal number. 2. T…

We have no idea what their contract is. But two business days without a reply isn’t exactly a long time. Especially if they are conducting their own investigation and reproduction steps.

> But two business days without a reply isn’t exactly a long time

What???? We have 4 hour boots on the ground support with Supermicro and that's a few thousand dollars a year lol.

That doesn't make any sense for a customer as big as CF.

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#203

Earlier quoted context omitted.

They aren't telling the facts as they know them. Cloudflare themselves say that the information in the article is "speculation" (the article literally uses that term). Publicly casting blame based on speculation isn't something you do to someone that you want to have a good working relationship with, no matter how much money you pay them.

That's not true. This is behaviour that would be enough for me to pull the plug working with this DC as this is more than unacceptable.

> if you want to have a good working relationship with

What are you disagreeing with OP ?

He is talking about how to behave if you continue the relationship not whether to continue it .

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#204
post #176
post #106

Earlier quoted context omitted.

In all fairness the rest of the article is about that

So why spend so much time trying to shift blame to the vendor? They could've just started the article with something like: > Due to circumstances beyond our control the DC lost all power. We are still working with our vendors to investigate the cause. While such a failure should not have been possible, our systems are supposed to tolerate a complete loss of a DC.

Because a small handful of decisions probably led to the Clickhouse and Kafka services still being non-redundant at the datacenter level, which added up to one mistake. But a small handful of mistakes were made by the vendor. Calling out each one of them was bound to take up more page space.

The ordering that they list the mistakes would be a fair point to make though, in my opinion. They hinted at a mistake they made in their summary, but don't actually tell us point blank what it was until they tell us all the mistakes that their vendor made. I'd argue that was either done to make us feel some empathy for Cloudflare as being victims of the vendor's mistakes, misleading us somewhat. Or it was done that way because it was genuinely embarrassing for the author to write and subconsciously they want us to feel some empathy for them anyway. Or some combination of the two. Either way, I'll grant that I would have preferred to hear what went wrong internally before hearing what went wrong externally.

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#205

Earlier quoted context omitted.

[flagged]

Sounds like chatgpt doesn't want your business and tuned thier cloudflare settings accordingly. Conveniently cloudflare is getting the blame, which is presumably part of what they're paying for.

Yep, it's easy to spot folks who have never configured Cloudflare's WAF when they suggest Cloudflare is blocking their browser of choice instead of the website itself.

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#206

Earlier quoted context omitted.

We have no idea what their contract is. But two business days without a reply isn’t exactly a long time. Especially if they are conducting their own investigation and reproduction steps.

> But two business days without a reply isn’t exactly a long time What???? We have 4 hour boots on the ground support with Supermicro and that's a few thousand dollars a year lol. That doesn't make any sense for a customer as big as CF.

My impression from reading the writeup is that CF did receive support and communication from Flexential during the event (although not as much communication as they would have liked), but hasn't received confirmation from Flexential about certain root cause analysis things that would be included in a post-mortem.

Two days without support communications would be a long time, but my original comment about the two day period is about the post-mortem. It's totally reasonable IMO for a company to take longer than two days to gather enough information to correctly communicate a post-mortem for an issue like this, and IMO its unreasonable for CF to try to shame Flexential for that.

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#207

While debatably unprofessional to blame your vendor, I found this read to be fascinating. I'm sure there are blog posts that detail how data centers work and fail but it's rare to get that cross over from a software engineering context. It puts into perspective what it takes for an average data center of this class to fail: power outage, generator failure, and then battery loss.

I think what it really does is emphasise how common it is for crap to hit the fan when things go wrong - even with the best laid plans.

The DC almost certainly advertises the redundant power supplies, generator backups and battery failover in order to get the customers. But probably doesn't do the legwork or spend the money to make those things truly reliable. It's a bit like having automated backups - but never testing them and discovering they're empty when they're really needed.

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#210

Earlier quoted context omitted.

Even "we don't know why our data center is failing, but we're sending a team over to physically investigate now" would have been A+ communication in the moment.

Everything was on the status page since the start? DC related updates: > Update - Power to Cloudflare’s core North America data center has been partially restored. Cloudflare has failed over some core services to a backup data center, which has partially remediated impact. Cloudflare is currently working to restore the remaining affected services and bring the core North America data center back online. Nov 02, 2023…

I've got no knock on the status page. Cloudflare is disappointed in the lack of notification from their data center provider, and Cloudflare customers are disappointed in the lack of notification from their service provider.

Instead of defending what was done and calling that good enough, Cloudflare should use this as an opportunity to commit to reevaluating the strategy for customer outreach during major service failures. If that's what Cloudflare expects from its service providers, that's what Cloudflare should provide to its customers.

Post reply on HN