Live data from Hacker News

Teaching a new way to prevent outages at Google

sre.google

11–20 of 44 posts

Re: Teaching a new way to prevent outages at Google

#11
post #5

This is peak corporate drivel—bloated storytelling, buzzwords everywhere, and a desperate attempt to make an old idea sound revolutionary. The article spends paragraphs on some childhood radio repair story before awkwardly linking it to STPA, a safety analysis method that’s been around for decades. Google didn’t invent it, but they act like adapting it for software is a major breakthrough. Most of the piece is just f…

I'm not sure if things have changed over the past five years, but this is exactly the stuff you'd throw in a promotion packet or maybe in a performance (perf) review to hit that mythical "superb" rating. The breaking point for me (and why I left after almost a decade) was when people started getting high ratings for fixing things they had an original hand in causing. Honestly, the comfiest job in the world if you're…

You can’t swoop in and be a hero and make impact without a meteor.

Re: Teaching a new way to prevent outages at Google

#12
post #7

I don't understand and I really really want to. This seems so cool at a scale that I can't fathom. Tell me specifically how it's done at google with regards to a specific service, at least enough information to understand what's going on. Make it concrete. Like "B lacks feedback from C", why is this bad? You've told me absolutely nothing and it makes me angry.

This has really always been the case with Google philosophy docs. They tend to be very abstract and academic.

The biggest danger is taking everything at face value and structuring your work or organization the same exact way based solely on these documents. The reality is, the vast majority of companies are not Google and will never encounter Google’s problems. That’s not where the value is though.

Re: Teaching a new way to prevent outages at Google

#13
post #5

This is peak corporate drivel—bloated storytelling, buzzwords everywhere, and a desperate attempt to make an old idea sound revolutionary. The article spends paragraphs on some childhood radio repair story before awkwardly linking it to STPA, a safety analysis method that’s been around for decades. Google didn’t invent it, but they act like adapting it for software is a major breakthrough. Most of the piece is just f…

I'm not sure if things have changed over the past five years, but this is exactly the stuff you'd throw in a promotion packet or maybe in a performance (perf) review to hit that mythical "superb" rating. The breaking point for me (and why I left after almost a decade) was when people started getting high ratings for fixing things they had an original hand in causing. Honestly, the comfiest job in the world if you're…

What I’ve been seeing from Google’s products lately suggests that these are the only ones still there. It’s a house of cards built with professional bullshitters. Google’s culture has entered or is already deep within the bullshit era.

Re: Teaching a new way to prevent outages at Google

#14

> "The class itself is very well structured. I've heard about STPA in past years, but this was the first time I saw it explained with concrete examples. The Google example at the end was also really helpful." But the article itself contains no concrete examples.

If you can like examples from outside Google, STPA seems to have been around for years:

https://kagi.com/search?q=STPA&r=no&sh=6ZXVCq1feUflSKjoBMMXm...

Re: Teaching a new way to prevent outages at Google

#15
Thanks to all the people here pointing out how bloated, overly broad and useless this is. I went to read it thinking I would pick up something applicable and it was written in such a overwrought humanless style that I gave up learning nothing and thought the problem was me. I am glad to learn I am not alone.

Re: Teaching a new way to prevent outages at Google

#16
STAMP/STPA work well as a model and methodology for complex systems, I was interested in them a while ago in the context of cyber risk quantification. Having a fairly easy model to reason about unsafe control action is not a given in other approaches. I just wish they were adopted by more companies, I have seen too many of them stuck with ERM-based frameworks that do no make sense most of the time when scaled down to working at the system level granularity.

Re: Teaching a new way to prevent outages at Google

#17
post #7

I don't understand and I really really want to. This seems so cool at a scale that I can't fathom. Tell me specifically how it's done at google with regards to a specific service, at least enough information to understand what's going on. Make it concrete. Like "B lacks feedback from C", why is this bad? You've told me absolutely nothing and it makes me angry.

This link at the bottom is less confusing:

https://www.usenix.org/publications/loginonline/evolution-sr...

Re: Teaching a new way to prevent outages at Google

#18
post #5

This is peak corporate drivel—bloated storytelling, buzzwords everywhere, and a desperate attempt to make an old idea sound revolutionary. The article spends paragraphs on some childhood radio repair story before awkwardly linking it to STPA, a safety analysis method that’s been around for decades. Google didn’t invent it, but they act like adapting it for software is a major breakthrough. Most of the piece is just f…

I'm not sure if things have changed over the past five years, but this is exactly the stuff you'd throw in a promotion packet or maybe in a performance (perf) review to hit that mythical "superb" rating. The breaking point for me (and why I left after almost a decade) was when people started getting high ratings for fixing things they had an original hand in causing. Honestly, the comfiest job in the world if you're…

By "had a hand in causing" do you mean "they should have prevented it", or do you just mean "they were involved in the causation"? Because sometimes you're forced to do things you know are wrong, because that's what other people are making you do, and in that case you still "have a hand" in causing.

Re: Teaching a new way to prevent outages at Google

#19
post #9

> In one particular case at Google, a software controller–acting on bad feedback from another software system–determined that it should issue an unsafe control action. It scheduled this action to happen after 30 days. Even though there were indicators that this unsafe action was going to occur, no software engineers–humans–were actually monitoring the indicators. So, after 30 days, the unsafe control action occurred,…

If you're referring to the time they nuked an Australian retirement fund's VMware setup, no, that was basically a billing screwup. An operator left a field blank, the system assumed that meant a 1-year expiry, and dutifully deleted it after 1 year was up.

https://cloud.google.com/blog/products/infrastructure/detail...

Re: Teaching a new way to prevent outages at Google

#20
post #5

This is peak corporate drivel—bloated storytelling, buzzwords everywhere, and a desperate attempt to make an old idea sound revolutionary. The article spends paragraphs on some childhood radio repair story before awkwardly linking it to STPA, a safety analysis method that’s been around for decades. Google didn’t invent it, but they act like adapting it for software is a major breakthrough. Most of the piece is just f…

Oh wow, shallow communication performative piece in a way ?
Post reply on HN