Live data from Hacker News

Preliminary Post Incident Review

crowdstrike.com

101–110 of 227 posts

Re: Preliminary Post Incident Review

#101
post #55

Earlier quoted context omitted.

> Architects likely do not have a choice. Architects don't have a choice, CTO are well paid to golf with the CEO and delegate to their teams, Auditors just audit but are not involved with the technical implementations, Developers just develop according to the Spec, and Security team just are a pain in the ass. Nobody owns it... Everybody get's well paid, and at the end we have to get lessons learned...It's a s*&^&t s…

Seems like everyone thinks that Execs play golf with another Execs to seal the deal regardless how b0rken the system is. That CTO's job is on the line if the system can't meet the requirement, more so if the system is fucked. To think that every CTO is dumbass is like saying "everyone is stupid, except me, of course"

Not all CTO...but you just saw hundreds of companies, who could do better....

Re: Preliminary Post Incident Review

#102
post #55

Earlier quoted context omitted.

> Architects likely do not have a choice. Architects don't have a choice, CTO are well paid to golf with the CEO and delegate to their teams, Auditors just audit but are not involved with the technical implementations, Developers just develop according to the Spec, and Security team just are a pain in the ass. Nobody owns it... Everybody get's well paid, and at the end we have to get lessons learned...It's a s*&^&t s…

Some industries are forced by regulation or liability to have something like crowdstrike deployed on their systems. And crowdstrike doesn't have a lot of alternatives that tick as many checkboxes and are as widely recognized.

Please give me an example of that specific regulation.

Re: Preliminary Post Incident Review

#103
post #99

Earlier quoted context omitted.

> After internal QA, you release to small internal dev team, then to select members of other depts willing to dog-food it, then limited external partners then GA What about AV definition update for 0day swimming in the tubes right now?

Sure, those have happened before, but nothing with an impact like last weekend. That's inexcusable. At least definitions can update themselves out of trouble.

What do you refer to "those have happened before"?

Isn't that what happened? Not a software update, not an AV-definition update but more so an AV-definition "data" update. At least that's how I interpret "Rapid Response Content"

Re: Preliminary Post Incident Review

#104

There’s only one sentence that matters: "Provide customers with greater control over the delivery of Rapid Response Content updates by allowing granular selection of when and where these updates are deployed." This is where they admit that: 1. They deployed changes to their software directly to customer production machines; 2. They didn’t allow their clients any opportunity to test those changes before they took effe…

Unfortunately, putting the onus on risk adverse organizations like hospitals and governments to validate the AV changes means they just won't get pushed and will be chronically exposed. That said, maybe Crowdstrike should considering validating every step of the delivery pipeline before pushing to customers.

> Unfortunately, putting the onus on risk adverse organizations like hospitals and governments to validate the AV changes means they just won't get pushed and will be chronically exposed.

I have a similar feeling.

At the very least perhaps have an "A" and a "B" update channel, where "B" is x hours behind A. This way if, in an HA configuration, one side goes down there's time to deal with it while your B-side is still up.

Re: Preliminary Post Incident Review

#105
post #42
post #31

> How Do We Prevent This From Happening Again? > Software Resiliency and Testing > * Improve Rapid Response Content testing by using testing types such as: > * Local developer testing So no one actually tested the changes before deploying?!

And why is it "local developer testing" and not CI/CD. This makes them look like absolute amateurs.

They don't care, CI/CD, like QA, is considered a cost center for some of these companies. The cheapest thing for them is to offload the burden of testing every configuration onto the developer, who is also going to be tasked with shipping as quickly as possible or getting canned.

Claw back executive pay, stock, and bonuses imo and you'll see funded QA and CI teams.

Re: Preliminary Post Incident Review

#107
post #81

Earlier quoted context omitted.

How do we know it hasn't?

If it happened, the industry would have known by now. The group behind it will come out to the public.

This would be the kind of vulnerability that would be worth millions of dollars and used for targeted attacks and/or by state actors. It could take years to uncover (like Pegasus, which took 5 years to be discovered) or never be uncovered at all.

Re: Preliminary Post Incident Review

#108

There’s only one sentence that matters: "Provide customers with greater control over the delivery of Rapid Response Content updates by allowing granular selection of when and where these updates are deployed." This is where they admit that: 1. They deployed changes to their software directly to customer production machines; 2. They didn’t allow their clients any opportunity to test those changes before they took effe…

> They deployed changes to their software directly to customer production machines; 2. They didn’t allow their clients any opportunity to test those changes before they took effect; and 3. This was cosmically stupid and they’re going to stop doing that.

Is it really all that surprising? This is basically their business model - its a fancy virus scanner that is supposed to instantly respond to threats.

Re: Preliminary Post Incident Review

#109
post #40

Earlier quoted context omitted.

* Further testing of the file was skipped because of "trust in the checks performed in the Content Validator" and successful tests of previous versions that's crazy. How costly can it be to test the file fully in a CI job? I fail to see how this wasn't implemented already.

> How costly can it be to test the file fully in a CI job? It didn't need a CI job. It just needed one person to actually boot and run a Windows instance with the Crowdstrike software installed: a smoke test. TFA is mostly an irrelevent discourse on the product architecture, stuffed with proprietary Crowdstrike jargon, with about a couple of paragraphs dedicated to the actual problem; and they don't mention the non-e…

They mentioned they do dogfooding. Wonder why it did not work for this update.

Re: Preliminary Post Incident Review

#110
post #70

1) Everything went mostly well 2) The things that did not fail went so great 3) Many many machines did not fail 4) macOS and Linux unaffected 5) Small lil bug in the content verifier 6) Please enjoy this $10 gift card 7) Every windows machine on earth bsod'd but many things worked

Fun post, but I'll state the obvious because I think many people do believe that every Windows machine BSOD'd. It was only ones with Crowdstrike software. Which is apparently very common but isn't actually pre-installed by Microsoft in Windows, or anything like that.

Source: work in a Windows shop and had a normal day.

Post reply on HN