Live data from Hacker News

Software engineering lessons from RCAs of greatest disasters

anoopdixith.com

121–130 of 150 posts

Re: Software engineering lessons from RCAs of greatest disasters

#121

Earlier quoted context omitted.

I would generally pass this comment by, but it's just so distastefully hostile because you totally missed the point. GP's comment was expressing sardonic disbelief that a modern jet wouldn't be able to receive remote software updates, considering it's so ubiquitous and reliable in other fields, even those with much, much lower costs. Not that developers don't release faults.

Imagine remotely bricking a fleet of fighter jets. https://news.ycombinator.com/item?id=35983866

> Imagine remotely bricking a fleet of fighter jets.

> https://news.ycombinator.com/item?id=35983866

That's about routers, was that the article you meant?

Re: Software engineering lessons from RCAs of greatest disasters

#122
post #73

Earlier quoted context omitted.

In terms of engineering (QC, processes etc), modern day software industry is worse than almost any other industry out there. :-( And no, just plain complexity or fast-moving environment, is a factor but not the issue. It's that steps are skipped which are not skipped in other branches of engineering (eg. continous improvement of processes, learning from mistakes & implementing those. In software land: same mistakes m…

> In terms of engineering (QC, processes etc), modern day software industry is worse than almost any other industry out there. :-( How do you know that?

Take airplane safety: plane crashes, cause of the crash is thoroughly investigated, report recommends procedures to avoid that type of cause for planecrashes. Sometimes such recommendations become enforced across the industry. Result: air travel safer & safer to the point where sitting in a (flying!) airplane all day is safer than sitting on a bench on the street.

Building regulations: similar.

Foodstuffs (hygiene requirements for manufacturers): similar.

Car parts: see ISO9000 standards & co.

Software: eg. memory leaks - been around forever, but every day new software is released that has 'm.

C: ancient, not memory safe, should really only be used for niche domains. Yet it still is everywhere.

New AAA game: pay $$ after year(s?) of development, download many-MB patch on day 1 because game is buggy. Could have been tested better, but released anyway 'cause getting it out & making sales weighed heavier than shipping reliable working product.

All of this = not improving methods.

I'm not arguing C v. Rust here or whatever. Just pointing out: better tools, better procedures exist, but using them is more exception than the rule.

Like I said the list goes on. Other branches of engineering don't (can't) work like that.

Re: Software engineering lessons from RCAs of greatest disasters

#123
post #73

Earlier quoted context omitted.

> In terms of engineering (QC, processes etc), modern day software industry is worse than almost any other industry out there. :-( How do you know that?

Take airplane safety: plane crashes, cause of the crash is thoroughly investigated, report recommends procedures to avoid that type of cause for planecrashes. Sometimes such recommendations become enforced across the industry. Result: air travel safer & safer to the point where sitting in a (flying!) airplane all day is safer than sitting on a bench on the street. Building regulations: similar. Foodstuffs (hygiene re…

Exactly. The driving force is there but what is also good is that the industry - for the most part at least - realizes that safety is what keeps them in business. So not only is there a structure of oversight and enforcement, there is also an strongly internalized culture of safety created over decades to build on. An engineer that would propose something obviously unsafe would not get to finish their proposal, let alone implement it.

In 'regular' software circles you can find the marketing department with full access to raw data and front end if you're unlucky.

Re: Software engineering lessons from RCAs of greatest disasters

#124
post #45

One of my favourites that I ever heard about was from a friend of mine who used to work on safety-critical systems in defense applications. He told me fighter jets have a safety system that disables the weapons systems if a (weight) load is detected on the landing gear so that if the plane is on the ground and the pilot bumps the wrong button they don't accidentally blow up their own airbase[1]. So anyway when the Eu…

> this cost several millions to redeploy the (one-line) fix to actually check the weight from the sensor was less than the threshold Well maybe this is the other, compounding problem. Engineering complex machines with such a high cost of bugfix deployment seems like a big issue. It's funny that as an industry we now know how to safely deploy software updates to hundreds of millions of phones, with security checks, si…

> Or maybe, a few millions is like only a thousand hours of flying in jet fuel costs alone, not a big deal...

Pretty much tbh. For example, the development of the Saab JAS 39 Gripen (JAS-projektet) is the most expensive industrial project in modern Swedish history at a cost of 120+ billion SEK (11+ billion USD).

It was also almost cancelled after a very public crash in central Stockholm at the 1993 Stockholm Water Festival [1]. A crash that should not have happened because the flight should not have been approved in the first place, because they weren't yet confident that they'd completely solved the Pilot-Induced Oscillation (PIO) related issues that wrecked the first prototype 4 years prior (with the same test pilot) [2].

It was basically a miracle that no one was killed or seriously hurt in the Stockholm crash, had the plane hit the nearby bridge or any of the other densely crowded areas then it would've been a very different story.

[1] https://youtu.be/mkgShfxTzmo?t=122

[2] https://www.youtube.com/watch?v=k6yVU_yYtEc

Re: Software engineering lessons from RCAs of greatest disasters

#125

One of my favourites that I ever heard about was from a friend of mine who used to work on safety-critical systems in defense applications. He told me fighter jets have a safety system that disables the weapons systems if a (weight) load is detected on the landing gear so that if the plane is on the ground and the pilot bumps the wrong button they don't accidentally blow up their own airbase[1]. So anyway when the Eu…

Regarding [1], was it Red Heat, True Lies, or Eraser?

Re: Software engineering lessons from RCAs of greatest disasters

#126

One of my favourites that I ever heard about was from a friend of mine who used to work on safety-critical systems in defense applications. He told me fighter jets have a safety system that disables the weapons systems if a (weight) load is detected on the landing gear so that if the plane is on the ground and the pilot bumps the wrong button they don't accidentally blow up their own airbase[1]. So anyway when the Eu…

Regarding [1], was it Red Heat, True Lies, or Eraser?

True Lies was an airborne Harrier, not a Russian plane on the ground. So, while there are many reasons the scene was unrealistic, “weight on gear disables weapons” isn’t one.

Re: Software engineering lessons from RCAs of greatest disasters

#127
post #103

Earlier quoted context omitted.

The Mark 14 ended-up being a really good torpedo by the end of WWII. It even remained in service until the 80ies. In truth, and going back to this subject, the Mark 14 debacle highlights the need for a good and unbiased QA. This also holds true for software engineering.

My understanding is the BeuOrd (or BeuShip? I don't remember which) "didn't want to waste money on testing it", so instead we wasted hundreds of them fired at japanese shipping that didn't even impact their target, or never had a hope of detonating. Remember these kind of things next time someone pushes for move fast and break things in the name of efficiency and speed. Slow is fast.

Pre-war, it was more a case of "Penny wise and Pound foolish" partly due to budget limitation (they did things like testing only with foam warheads to recover test torpedoes).

But after Perl Harbor, a somewhat biased BuOrd was reluctant to admit the mark 14 flaws. It took a few "unauthorized" tests and 2 years to fix the issues.

In fairness, this sure makes for an entertaining story (ex Drachinifel video on yt), but I'm not completely sold on the depiction of BuOrd as some sort of arrogant bureaucrats. However, bias and pride (plus other issues like low production) certainly have played a role in the early mark 14 debacle.

Going back to software development, I'm always amazed how bugs immediately pop-up whenever I put a piece of software in the hands of users for the first time, and that's regardless how well I tested it. I try to be as thorough as possible, but being the developer I'm always bias, often tunnel visioning on one way to use the software I created. That's why, in my opinion you need some form of external QA/testing (like these "unauthorized" Mark 14 tests).

Re: Software engineering lessons from RCAs of greatest disasters

#128

Earlier quoted context omitted.

Pretty epic. I was working for a webhosting company, and someone asked me to rush a change just before leaving. Instead of updating 1500 A records, I updated about 50k. Someone senior managed to turn off the cron though, so what I actually lost was the delta of changes between last backup and my SQL. I was in the room for this though: https://www.theregister.com/2008/08/28/flexiscale_outage/

I love the title to that article "Engineer accidentally deletes cloud". It's like a single individual managed to delete the monolithic cloud where everyone's files are stored.

Bare in mind, this was a small startup in 2008 that claims to be the 2nd cloud in the world ( read on-demand iaas provider ).

Flexiscale at the time was a single region backed by a netapp. Each VM essentially had a thin-provisioned lun ( logical volume ), basically you copy on write the underlying OS image.

So when someone accidently deletes vol0, they take out a whopping 6TB of data, that takes a ~20TB to restore because you're rebuilding filesystems from safe mode ( thanks netapp support ). It's fairly monolithic in that sense.

I guess I was 23 at the time, but I'd written the v2 API, orchestrator & scheduling later. It was fairly naive, but filled the criteria of a cloud, i.e. elastic, on-demand, metered usage, despite using a SAN.

Re: Software engineering lessons from RCAs of greatest disasters

#129

Earlier quoted context omitted.

Regarding [1], was it Red Heat, True Lies, or Eraser?

True Lies was an airborne Harrier, not a Russian plane on the ground. So, while there are many reasons the scene was unrealistic, “weight on gear disables weapons” isn’t one.

I know, such a great movie, total classic of my childhood! It was the only R-rated film my mother ever allowed and even endorsed us watching, "Because Jamie Lee Curtis is hot." :D

I wasn't sure if there might've been a scene I'd forgotten.

Re: Software engineering lessons from RCAs of greatest disasters

#130

Earlier quoted context omitted.

I love the title to that article "Engineer accidentally deletes cloud". It's like a single individual managed to delete the monolithic cloud where everyone's files are stored.

That is eerily similar to what happened to us in IBM "Cloud", in a previous gig. An engineer was doing "account cleanup" and somehow our account got on the list and all our resources were blown away. The most interesting conversation was convincing the support person, that those deletion audit events were in fact not us, but rather (according to the engineer's Linked-In page) an SRE at IBM.

This was ~14 years ago and both MS & AWS had loss of data incidents iirc.
Post reply on HN