Live data from Hacker News

Post Mortem on Cloudflare Control Plane and Analytics Outage

blog.cloudflare.com

221–230 of 241 posts

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#221
post #179

Its weird that upon reading this post, I have less confidence in Cloudflare. They basically browbeat Flexential for behaving unprofessional, which, yes, they probably did. However the fact that this causes entire systems that people rely on to go down is a massive redundancy failure on Cloudflares part, you should be able to nuke one of these datacentres and still maintain services. Very worrying is they start by sta…

I'm also a bit surprised about Hillsboro. The FEMA is assuming that when (not if) The Big One hits, everything west of I-5 is going to be toast. Is placing the entirety of such a critical cluster in a known earthquake and tsunami zone a good idea? It looks like their disaster recovery to Europe didn't really work either...

Yeah. Moreover, looking at the map, the DCs around Hillsboro are terrifyingly close to each other.

By the way, assuming an ideal control plane (in contrast to data plane) would be 3 DCs at a distance of about 20-40 miles, are there any mitigation techniques so that a seismic event which destroys a single DC doesn't also sever the comms between the remaining two?

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#222

Earlier quoted context omitted.

If Flexential and PGE aren't sharing information or otherwise cooperating as much as Cloudflare might like, then going public with some speculation might be an attempt at applying some pressure to get to the bottom of what happened. It might also be an effort to get out in front of the story before someone else does the speculating. In any case, with at least three parties involved, with multiple interconnected syste…

>If Flexential and PGE aren't sharing information or otherwise cooperating as much as Cloudflare might like, then going public with some speculation might be an attempt at applying some pressure to get to the bottom of what happened. It's been 2 days. I doubt PGE or Flexential even have root caused it yet, and even if they have, good communication takes time. You don't throw someone under the bus and smear their name…

[flagged]

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#223
post #48

Earlier quoted context omitted.

> I find that pretty embarrassing for a company like Cloudflare, which powers such relevant parts of the internet. Bah, who cares about such unimportant details, what's important is that ~dev velocity~ was reaaally high right until that moment! > We were also far too lax about requiring new products and their associated databases to integrate with the high availability cluster. Cloudflare allows multiple teams to inn…

I’m going to leave out some details but there was a period of time where you could bypass cloudflare’s IP whitelisting by using Apple’s iCloud relay service. This was fixed but to my knowledge never disclosed.

Saw.t

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#224
post #176
post #106

Earlier quoted context omitted.

In all fairness the rest of the article is about that

So why spend so much time trying to shift blame to the vendor? They could've just started the article with something like: > Due to circumstances beyond our control the DC lost all power. We are still working with our vendors to investigate the cause. While such a failure should not have been possible, our systems are supposed to tolerate a complete loss of a DC.

I don't think I read it as charged as you did

Here's what happened, here's what went wrong, here's what we did wrong, here's our plans to avoid it happening again

Seems like a standard post mortem tbh

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#225

Earlier quoted context omitted.

I've got no knock on the status page. Cloudflare is disappointed in the lack of notification from their data center provider, and Cloudflare customers are disappointed in the lack of notification from their service provider. Instead of defending what was done and calling that good enough, Cloudflare should use this as an opportunity to commit to reevaluating the strategy for customer outreach during major service fai…

? You want Cloudflare to update every customer for an issue that they probably aren't affected with ( except when changing things) ? Who even does that when you've got so many customers? That's exactly what why the status page is there: https://www.cloudflarestatus.com/ The DC obviously didn't have any means to update their customers.

I don't want that. Cloudflare's customers want that. Cloudflare was embarrassed and needs to listen to the feedback they're receiving.

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#226
post #106

Earlier quoted context omitted.

In all fairness the rest of the article is about that

Slightly less than half, and the bottom half, so that people just skimming over it will mostly remember the DC operators' problems, not Cloudflare's own. This is very deliberately manipulative.

It is of course possible they've shuffled things around since this was posted but it seems that the first part addresses their system failings.

5th paragraph to the 9th are Cloudflare's "we buggered up" before they get to the power segment. They then continue with the "this is our fault for not being fully HA" after the power bit.

Each to their own, I'm going to read it as a regular old post mortem on this one.

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#227
You can run ClickHouse cluster across multiple datacenters. It will survive the failure of a single datacenter while being available for writes and reads, and the failure of two out of three datacenters while being available for reads. It works well when RTT between datacenters is less than 30 ms. If they are more distant, it will still work, but you will notice a quite high latency on INSERTs due to the distributed consensus.

I've run a ClickHouse cluster with hundreds of bare-metal machines distributed across three datacenters in two countries at my previous job. It survived power failures (multiple), a flood (once), and network connectivity issues (regular). This cluster was used for logging and analytics :)

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#228
Why would any supplier want to do business with Cloudflare now? You have 1.8MW of datacenter space to lease, you have a few interested parties, how could you not view Cloudflare as a huge reputational risk? Why even do the business? Why not lease that space to someone else? Moreover, why renew the existing deals with Cloudflare?

Does Cloudflare have a plan to move 200+ racks in Oregon if that supplier decides just not to renew that deal whenever it comes up next? Are Cloudflare claiming they were able to build a technical plan which gets their architecture away from this site being a SPOF before the deal is up to renew, or is the CEO making a gamble here again?

Cloudflare have demonstrated their willingness to create reputational issues for suppliers by publicly shaming two of them recently, and here, in only about 2 days from incident. One interpretation of this blog would be Cloudflare are a very unreasonable customer and one who is willing to post incomplete or informal information from their suppliers. Cloudflare also chose to focus the first half of a lengthy postmortem on blaming the supplier and only then on their own culpability for the outage, despite it clearly being a shared responsibility.

One of the diagrams Cloudflare have posted is clearly marked "Proprietary and Confidential". Do Cloudflare have permission to post that? It's not clearly stated that they do. Should other suppliers expect when the sh*t hits the fan that any sensitive information they've shared will be part of a blog?

Most of the "Lessons and Remediation" section is stuff Cloudflare could have worked on at any point in advance of a major incident, and Cloudflare's senior management have quite clearly chosen not to prioritize that work until today, when forced to by this major incident.

When signing large deals, Cloudflare will frequently have to complete 'Supplier Disclosures', and they are also making claims through industry-standard certifications [1] like ISO, SOC and FedRAMP. Most of those will ask questions about the disaster recovery and business continuity plans and Cloudflare will have (repeatedly) attested they are adequate, something that this blog clearly demonstrates was a misrepresentation of their true capabilities.

Will there be an SEC disclosure coming out of this considering it could have material impacts on the business, which is publicly traded? Was there any requirement that the SEC disclosure come first, or be concurrent with a blog?

[1] https://www.cloudflare.com/trust-hub/compliance-resources/

Re: Post Mortem on Cloudflare Control Plane and Analytics Outage

#229

Earlier quoted context omitted.

? You want Cloudflare to update every customer for an issue that they probably aren't affected with ( except when changing things) ? Who even does that when you've got so many customers? That's exactly what why the status page is there: https://www.cloudflarestatus.com/ The DC obviously didn't have any means to update their customers.

I don't want that. Cloudflare's customers want that. Cloudflare was embarrassed and needs to listen to the feedback they're receiving.

There is literally not a single cloud company doing that.

Even those that had complete outages.

Post reply on HN