Earlier quoted context omitted.
“Decided our fate in a microsecond.”
It would hide out and subtly distort our culture, slowly driving the society mad, and slowly driving us all mad...for the lulz!
Google Cloud Networking Incident Postmortem
101–110 of 190 posts
Re: Google Cloud Networking Incident Postmortem
#102Re: Google Cloud Networking Incident Postmortem
#103I was curious to know how cascading failures in one region effected other regions. Impact was " ...increased latency, intermittent errors, and connectivity loss to instances in us-central1, us-east1, us-east4, us-west2, northamerica-northeast1, and southamerica-east1." Answer, and the root cause summarized: Maintenance started in a physical location, and then "... the automation software created a list of jobs to des…
> Debugging the problem was significantly hampered by failure of tools competing over use of the now-congested network. Man that's got to suck.
A very simple example, you do something stupid on a remote machine (either high network usage or CPU usage) over SSH then you can't undo it because SSH becomes unresponsive
Re: Google Cloud Networking Incident Postmortem
#104My burning question is what is a "relatively rare maintenance event type"?
You normally prepare for such a task for a month, and then you hope it will work. In my case (I brought down one the core DNS in Austria for a few minutes, for a very trivial oversight) everyone knew, and after the caches ran out we immediately restored the backup. We weren't on page one in the news as Google.
In the Google case they had no idea of the root cause, so they had to run after this guy who caused it. Only after 4 hours they found him, and they could stop this job. Reminds me a bit of Chernobyl, where nobody told anybody.
Re: Google Cloud Networking Incident Postmortem
#105Having only ever seen one major outage event in person (at a financial institution that hadn't yet come up with an incident response plan; cue three days of madness), I would love to be a fly on the wall at Google or other well-established engineering orgs when something like this goes down. I'd love to see the red binders come down off the shelf, people organize into incident response groups, and watch as a root cau…
You might be interested in https://response.pagerduty.com/ , PagerDuty's major incident response process documentation - a good starting point for that red binder. Having been in the ringmasters seat for major incidents ranging from "relatively routine" to "it's all on fire", and had a ringside seat for a cloud provider outage of comparable magnitude to this one - it still fascinates me how creative solutions can get…
Re: Google Cloud Networking Incident Postmortem
#106Earlier quoted context omitted.
it's pretty boring. real life computers aren't at all like hackers or csi:cyber. except for the skateboards, all real sysadmins ride skateboards.
Documentary that includes all the technical details wouldn't be. Kind of like the ambulance shows that seem to be popular now, but more technical. Of course the target audience is probably tiny.
Re: Google Cloud Networking Incident Postmortem
#107I was curious to know how cascading failures in one region effected other regions. Impact was " ...increased latency, intermittent errors, and connectivity loss to instances in us-central1, us-east1, us-east4, us-west2, northamerica-northeast1, and southamerica-east1." Answer, and the root cause summarized: Maintenance started in a physical location, and then "... the automation software created a list of jobs to des…
> Debugging the problem was significantly hampered by failure of tools competing over use of the now-congested network. Man that's got to suck.
Re: Google Cloud Networking Incident Postmortem
#108Earlier quoted context omitted.
Yes this is also how it's done at other large orgs. But one key to a quick response is for every low-level team to have at least one engineer on call at any given time. This makes it so any SRE team can engage with true "owners" of the offending code ASAP. Also during an incident, fingers are never publicly/embarrassingly pointed nor are people blamed. It's all about identifying and fixing the issue as fast as possib…
I have mixed feelings about the finger pointing/public embarrassment thing. Usually the SRE is matured enough cause they have to be, however the individual teams might not be the same when it comes to reacting/handling the Incident report/postmortem. On a slightly different note, "low-level team to have at least one engineer on call at any given time" - this line itself is so true and at the same time it has so many…
There is an understanding of how it impacts your time, your energy and your life that is impressive? To be honest, I feel bad for being so macho about oncall at the org I ran and just having the leads take it all upon ourselves.
Re: Google Cloud Networking Incident Postmortem
#109I was curious to know how cascading failures in one region effected other regions. Impact was " ...increased latency, intermittent errors, and connectivity loss to instances in us-central1, us-east1, us-east4, us-west2, northamerica-northeast1, and southamerica-east1." Answer, and the root cause summarized: Maintenance started in a physical location, and then "... the automation software created a list of jobs to des…
> Debugging the problem was significantly hampered by failure of tools competing over use of the now-congested network. Man that's got to suck.
Re: Google Cloud Networking Incident Postmortem
#110“No, comrade. You’re mistaken. RBMK reactors don’t just explode.”