Live data from Hacker News

Preliminary Post Incident Review

crowdstrike.com

201–210 of 227 posts

Re: Preliminary Post Incident Review

#201

There’s only one sentence that matters: "Provide customers with greater control over the delivery of Rapid Response Content updates by allowing granular selection of when and where these updates are deployed." This is where they admit that: 1. They deployed changes to their software directly to customer production machines; 2. They didn’t allow their clients any opportunity to test those changes before they took effe…

Combined with this, presented as a change they could potentially make, it's a killer: > Implement a staggered deployment strategy for Rapid Response Content in which updates are gradually deployed to larger portions of the sensor base, starting with a canary deployment. They weren't doing any test deployments at all before blasting the world with an update? Reckless.

> our staging environment, which consists of a variety of operating systems and workloads

they have a staging environment at least, but no idea what they were running in it or what testing was done there.

Re: Preliminary Post Incident Review

#202
post #194

Earlier quoted context omitted.

I don’t have much sympathy for CrowdStrike but deploying slowly seems mutually exclusive to protecting against emerging threats. They have to strike a balance.

Even a staged rollout over a few hours would have made a huge difference here. "Slow" in the context of a rollout can still be pretty fast.

But it can also still be way too slow in the context of an exploit that is being abused globally.

Re: Preliminary Post Incident Review

#203

I work on a piece of software that is installed on a very large number of servers we do not own. The crowd strike incident is exactly our nightmare scenario. We are extremely cautious about updates, we roll it out very slowly with tons of metrics and automatic rollbacks. I’ve told my manager to bookmark articles about the crowdstrike incident and share it with anyone who complains about how slow the update process is…

> let [...] owners control when to update

The only acceptable update strategy for all software regardless of size or importance

Re: Preliminary Post Incident Review

#204

I work on a piece of software that is installed on a very large number of servers we do not own. The crowd strike incident is exactly our nightmare scenario. We are extremely cautious about updates, we roll it out very slowly with tons of metrics and automatic rollbacks. I’ve told my manager to bookmark articles about the crowdstrike incident and share it with anyone who complains about how slow the update process is…

I don’t have much sympathy for CrowdStrike but deploying slowly seems mutually exclusive to protecting against emerging threats. They have to strike a balance.

They need a lab full of canaries.

Re: Preliminary Post Incident Review

#205

Earlier quoted context omitted.

I don’t have much sympathy for CrowdStrike but deploying slowly seems mutually exclusive to protecting against emerging threats. They have to strike a balance.

They need a lab full of canaries.

yeah, I don't get the "we couldn't have tested it" crap, because "something happens to the payload after we tested it". Create a fake downstream company and put a bunch of machines in it. That's your final test before releasing to the rest of the world.

Re: Preliminary Post Incident Review

#206
post #194

Earlier quoted context omitted.

I don’t have much sympathy for CrowdStrike but deploying slowly seems mutually exclusive to protecting against emerging threats. They have to strike a balance.

Even a staged rollout over a few hours would have made a huge difference here. "Slow" in the context of a rollout can still be pretty fast.

Sure but GP is praising "deploy so slowly that people complain."

Re: Preliminary Post Incident Review

#207

Earlier quoted context omitted.

Why can't they just do it more like Microsoft security patches, making them mandatory but giving admins control over when they're deployed?

That would be equivalent to asking "would you prefer your fleet to bluescreen now, or later" in this case.

Those eager would take it immediately, those conservative would wait (and be celebrated by C-suite later when SHTF). Still a much better scenario than what happened.

Re: Preliminary Post Incident Review

#208

Earlier quoted context omitted.

no, it's one of most well written PIR's I've seen. It establishes terms and procedures after communicating that this isn't an RCA, then they detail the timeline of tests and deployments done and what went wrong. They were not excessively verbose or terse. This is the right way of communicating to the intended audience. It is both technical people, executives and law makers alike that will be reading this. They commun…

If you think this is good, go look at a Cloudflare postmortem. The fly.io ones are good too. Way less obscure language, way more detail and depth, actually owning the mistakes rather than vaguely waffling on. This write up from CrowdStrike is close to being functionally junk.

One of the first things they've stated is that this isn't an RCA (deep dive analysis) like cloudflare and fly.io's, that's not what this is. This is to brief customers and the public of their immediate post-mortem understanding of what happened. The standard for that is different than an RCA.

Re: Preliminary Post Incident Review

#209
post #148

Earlier quoted context omitted.

> They didn’t allow their clients any opportunity to test those changes before they took effect I’d argue that anyone that agrees to this is the idiot. Sure they have blame for being the source of the problem, but any CXO that signed off on software that a third party can update whenever they’d like is also at fault. It’s not an “if” situation, it’s a “when”.

I felt exactly the same when I read about the outage. What kind of CTO would allow 3rd party "security" software to automatically update? That's just crazy. Of course, your own security team would do some careful (canary-like) upgrades locally... run for a bit... run some tests, then sign-off. Then upgrade in a staged manner.

Pretty sure many people see the point of having Falcon as a reason to not have an internal security team.

Outsource everything.

Re: Preliminary Post Incident Review

#210
post #182

"Based on the testing performed before the initial deployment of the Template Type (on March 05, 2024), trust in the checks performed in the Content Validator, and previous successful IPC Template Instance deployments, these instances were deployed into production." So they did not test this update at all, even locally. Its going to be interesting how this plays out in courts. The contract they have with us limits th…

As I understand, it is incredibly difficult to prove "gross negligence". It is better to pressure them to settle in a giant class action lawsuit. I am curious what the total amount of settlements / fines will be in the end. I guess ~2B USD.

Same here. Our losses were quite significant - between lost productivity, inability to provide services, inability of our clients to actually use contracted services, and having to fix their mess - its very easily in the millions.

And then there will be the costs of litigation. It was crazy in the IT department over the weekend, but not much less crazy in our legal teams, who were being bombarded with pitches from law firms offering help in recovery. It will be a fun space to watch, and this 'we haven't tested because we, like, did that before and nothing bad happened' statement in the initial report will be quoted in many lawsuits.

Post reply on HN