Live data from Hacker News

Software engineering lessons from RCAs of greatest disasters

anoopdixith.com

51–60 of 150 posts

Re: Software engineering lessons from RCAs of greatest disasters

#51
post #11

I think the software industry itself has accumulated enough bugs over the past few decades. E.g. F-22 navigation system core dumps when crossing the international date line: https://medium.com/alfonsofuggetta-it/software-bug-halts-f-2... Loss of Mars probe due to metric-imperial conversion error I've a few of these myself (e.g. a misplaced decimal that made $12mil into $120mil), but sadly cannot devulge details.

Does anyone have a similar compendium specifically for software engineering disasters?

Not of nasty bugs like the F-22 -- those are fun stories, but they don't really illustrate the systemic failures that led to the bug being deployed in the first place. Much more interested in systemic cultural/practice/process factors that led to a disaster.

Re: Software engineering lessons from RCAs of greatest disasters

#52
post #39

Earlier quoted context omitted.

> Every disaster is preventable. No, there is such a thing as residual risk and there are always disasters that you can't prevent such as natural disasters. But even then you can have risk mitigation and strategies for dealing with the aftermath of an incident to limit the effects. > Everything on the list was happening in human-engineered environments - as do most things that affect humans. That is precisely why the…

It is 2023. The damage of natural disasters can be mitigated. When the San Andreas fault goes it'll probably get an entry on that list with a "why did we build so much infrastructure on this thing? Why didn't we prepare more for the inevitable?". And this article is throwing out generic all-weather good sounding platitudes which are tangential to the disasters listed. He drew a comparison between the Challenger disas…

> It is 2023.

So? Mistakes are still being made, every day. Nothing has changed since the stone age except for our ability - and hopefully willingness - to learn from previous mistakes. If we want to.

> The damage of natural disasters can be mitigated.

You wish.

> When the San Andreas fault goes it'll probably get an entry on that list with a "why did we build so much infrastructure on this thing? Why didn't we prepare more for the inevitable?".

Excellent questions. And in fairness to the people living on the San Andreas fault - and near volcanoes, in hurricane alley and in countries below sea level - we have an uncanny ability to ignore history.

> And this article is throwing out generic all-weather good sounding platitudes which are tangential to the disasters listed.

I see these errors all the time in the software world, I don't care what hook he uses to again bring them to attention but they are probably responsible for a very large fraction of all software problems.

> He drew a comparison between the Challenger disaster and bitrot!

So let's see your article on this subject then that will obviously do a much better job.

> Anyone who thinks that is a profound connection should avoid the role of software architect.

Do you care? It would be better to say that those that fail to be willing to learn from the mistakes of others should avoid the role of software architect because on balance that's where the problems come from. You seem to have a very narrow viewpoint here: that because you don't like the precision or the links that are being made that you can't appreciate the intent and the subject matter. Of course a better article could have been written and of course you are able to dismiss it entirely because of its perceived shortcomings. But that is exactly the attitude that leads to a lot of software problems: the inability to ingest information when it isn't presented in the recipients preferred form. This throws out the baby with the bath water, the authors intent is to educate you and others on the ways in which software systems break and uses something called a narrative hook to serve as a framework. That these won't match 100% is a given. Spurious connection or not, documentation and actual fact creeping out of spec aka the normalization of deviation in disguise is exactly the lesson from the Challenger disaster and if you don't like the wording I'm looking forward to your improved version.

> Challenger was about catastrophic management and safety practices.

That was a small but critical part in the whole, I highly recommend reading the entire report on the subject, it makes for fascinating reading, there are a great many lessons to be learned from this.

https://www.govinfo.gov/content/pkg/GPO-CRPT-99hrpt1016/pdf/...

https://en.wikipedia.org/wiki/Rogers_Commission_Report

And many useful and interesting supporting documents.

> I mean, if we want to learn from Douglas Adams he suggested that we can deduce the nature of all things by studying cupcakes.

That's a complete nonsensical statement. Have you considered that your initial response to the article precludes you from getting any value from it?

> It is not useful to connect random things in other fields to random things in software.

But they are not random things. The normalization of deviation in whatever guise it comes is the root cause of many, many real world incidents, both in software as well as outside of it. You could argue with the wording, but not with the intent or the connection.

> Although I do appreciate the effort the gentleman went to, it is a nice site and the disasters are interesting. Just not relevantly linked to software in a meaningful way.

To you. But they are.

> > We are tied 1:1 to the fate of our star and may well go down with it > I'm just going to claim that is false and live in the smug comfort that when circumstances someday prove you right neither of us will be around to argue about it.

So, you are effectively saying that you persist in being wrong simply because the timescale works to your advantage?

> And if you can draw lessons from that which apply to practical software development then that is quite impressive.

Well, for starters I would argue that many software developers indeed create work that serves just long enough to hold until they've left the company and that that attitude is an excellent thing to lose and a valuable lesson to draw from this discussion.

Re: Software engineering lessons from RCAs of greatest disasters

#53
post #51
post #11

I think the software industry itself has accumulated enough bugs over the past few decades. E.g. F-22 navigation system core dumps when crossing the international date line: https://medium.com/alfonsofuggetta-it/software-bug-halts-f-2... Loss of Mars probe due to metric-imperial conversion error I've a few of these myself (e.g. a misplaced decimal that made $12mil into $120mil), but sadly cannot devulge details.

Does anyone have a similar compendium specifically for software engineering disasters? Not of nasty bugs like the F-22 -- those are fun stories, but they don't really illustrate the systemic failures that led to the bug being deployed in the first place. Much more interested in systemic cultural/practice/process factors that led to a disaster.

Yes, the RISK mailing list.

Re: Software engineering lessons from RCAs of greatest disasters

#54
post #44

As I heard one engineering leader say, "it's okay to make mistakes — once ". Meaning, we're all fallible, mistakes happen, but failure to learn from past mistakes is not optional. That said, a challenge I have frequently run into, and I feel is not uncommon, is a tension between the desire not to repeat mistakes and ambitions that do generally involve some amount of risk-taking. The former can turn into a fixation on…

Failure is not optional. Definitely true :)

Re: Software engineering lessons from RCAs of greatest disasters

#55
post #39

Earlier quoted context omitted.

> Every disaster is preventable. No, there is such a thing as residual risk and there are always disasters that you can't prevent such as natural disasters. But even then you can have risk mitigation and strategies for dealing with the aftermath of an incident to limit the effects. > Everything on the list was happening in human-engineered environments - as do most things that affect humans. That is precisely why the…

It is 2023. The damage of natural disasters can be mitigated. When the San Andreas fault goes it'll probably get an entry on that list with a "why did we build so much infrastructure on this thing? Why didn't we prepare more for the inevitable?". And this article is throwing out generic all-weather good sounding platitudes which are tangential to the disasters listed. He drew a comparison between the Challenger disas…

An Ariane 5 failed because of bitrot, so the headline comparison of rocket failures makes sense. Not testing software with new performance parameters before launch sounds like catastrophic management to me.

Re: Software engineering lessons from RCAs of greatest disasters

#57

One of my favourites that I ever heard about was from a friend of mine who used to work on safety-critical systems in defense applications. He told me fighter jets have a safety system that disables the weapons systems if a (weight) load is detected on the landing gear so that if the plane is on the ground and the pilot bumps the wrong button they don't accidentally blow up their own airbase[1]. So anyway when the Eu…

Torpedoes typically have an inertial switch which disarms them if they turn 180 degrees, so they don't accidentally hit their source. When a torpedo accidentally arms and activates on board a submarine (hot running) the emergency procedure is to immediately turn the sub around 180 degrees to disarm the torpedo.

Re: Software engineering lessons from RCAs of greatest disasters

#58
post #39

Earlier quoted context omitted.

It is 2023. The damage of natural disasters can be mitigated. When the San Andreas fault goes it'll probably get an entry on that list with a "why did we build so much infrastructure on this thing? Why didn't we prepare more for the inevitable?". And this article is throwing out generic all-weather good sounding platitudes which are tangential to the disasters listed. He drew a comparison between the Challenger disas…

> It is 2023. So? Mistakes are still being made, every day. Nothing has changed since the stone age except for our ability - and hopefully willingness - to learn from previous mistakes. If we want to. > The damage of natural disasters can be mitigated. You wish. > When the San Andreas fault goes it'll probably get an entry on that list with a "why did we build so much infrastructure on this thing? Why didn't we prepa…

So the article had a list of disasters and some useful lessons learned in its left and center columns. It also had lists of truisms about software engineering in the right column. They had nothing fundamental to do with each other.

For instance, it tries to draw an equivalence between "Titanic's Captain Edward Smith had shown an "indifference to danger [that] was one of the direct and contributing causes of this unnecessary tragedy." and "Leading during the time of a software crisis (think production database dropped, security vulnerability found, system-wide failures etc.) requires a leader who can stay calm and composed, yet think quickly and ACT." which are completely unrelated: one is a statement about needing to evaluate risks to avoid incidents, another is talking about the type of leadership needed once an incident has already happened. Similarly, the discussion about Chernobyl is also confused: the primary lessons there are about operational hygiene, but the article draws "conclusions" about software testing which is in a completely different lifecycle phase.

There are certainly lessons to be learned from past incidents both software and not, but the article linked is a poor place to do so.

Re: Software engineering lessons from RCAs of greatest disasters

#59
post #45

One of my favourites that I ever heard about was from a friend of mine who used to work on safety-critical systems in defense applications. He told me fighter jets have a safety system that disables the weapons systems if a (weight) load is detected on the landing gear so that if the plane is on the ground and the pilot bumps the wrong button they don't accidentally blow up their own airbase[1]. So anyway when the Eu…

> this cost several millions to redeploy the (one-line) fix to actually check the weight from the sensor was less than the threshold Well maybe this is the other, compounding problem. Engineering complex machines with such a high cost of bugfix deployment seems like a big issue. It's funny that as an industry we now know how to safely deploy software updates to hundreds of millions of phones, with security checks, si…

This makes no sense and is difficult to even respond to coherently.

> It's funny that as an industry we now know how to safely deploy software updates to hundreds of millions of phones, with security checks, signed firmwares, etc,

Either you're completely wrong, because we "as an industry" still push bugs and security flaws, or you're comparing two completely different things.

> doing that on applications with a super high unit price tag seems out of reach...

is true because of

> a few millions is like only a thousand hours of flying in jet fuel costs alone

like do you really think they spent millions pushing a line of code? or do you think it's just inherently expensive to fly a jet, and so doing it twice costs more?

Re: Software engineering lessons from RCAs of greatest disasters

#60
post #34

This is why software engineering is a protected profession in some parts of the world (Canada at least), as civil responsibility and safety, along with formal legal liability is part of licensure

Care to elaborate? I know professional engineers in Canada get a designation but I’m not aware of anything similar for software engineers.

Software engineers are the same as all other engineering professions, and regulated by the same provincial PEG associations. While most employers don't care about it, some software positions where the safety of people is in line (eg aeronautics) or there's a special stake do have requirements to employ professional software engineers.

I think you're actually not even supposed to call yourself an engineer unless you're a professional engineer.

Post reply on HN