Live data from Hacker News

Post-Mortem for Google Compute Engine’s Global Outage on April 11

status.cloud.google.com

51–60 of 368 posts

Re: Post-Mortem for Google Compute Engine’s Global Outage on April 11

#52

Earlier quoted context omitted.

> This is also why I am of afraid of a self driving cars and other such life critical software. There are going to be weird edge cases, what prevents you from reaching them? Yes. However, the current failure rate of human drivers being improved on is the standard I care about. http://www.cnbc.com/2015/10/29/crash-data-for-self-driving-c... > After crunching the data, Schoettle and Sivak concluded there's an average o…

So crashes like that include drunk drivers, drugged drivers, and texting drivers. What is the crashes per miles for a paying attention driver? If it is 1 per million miles, the self driving car would need to be a lot lower. Now if it was 4am and I am falling asleep at the wheel, I bet any self driving car would beat me. So cool to turn on, but maybe not for a daytime cruise...

Why are you comparing self-driving cars to exclusively a "paying attention driver"?

For self-driving cars to be safer than human drivers, there is no requirement that the self-driving cars should be better/safer than the best human driver... the self-driving car simply needs to be safer than the majority of humans.

Re: Post-Mortem for Google Compute Engine’s Global Outage on April 11

#53

This is very interesting. From the little I understand (sorry for using AWS terms as I am more versed with AWS than GCE) this can happen to AWS as well right? even if your software is deployed to multiple AZs / multiple regions, if bad routing / network configuration makes it through the various protection mechanisms then basically no amount of redundancy can help if your service is part of the non functional IP bloc…

I'm not sure the AWS network follows the same setup, AWS has very distinct blocks between the US/EU/APAC compared to GCP where you can inherit the same IP if you quickly delete/recreate instances in different regions?

Re: Post-Mortem for Google Compute Engine’s Global Outage on April 11

#54
post #42

Nice post mortem. That outtage gives GCE at best a four 9's reliability for 2016.

Based on the higher level status page: https://status.cloud.google.com/summary It looks like GCE uptime is well below four 9's reliability for a sliding 1 year timeframe.

Traynor was quoted in a networkworld article last year saying they aim for three and a half nines (99.95%). But you need to read into the incidents more carefully -- figuring out actual "uptime" is quite hard. Consider the longest-lasting incident:

  "On Tuesday 23 February 2016, for a duration of 
   10 hours and 6 minutes, 7.8% of Google Compute Engine
   projects had reduced quotas.  ...  Any resources that
   were already created were unaffected by this issue."
I'm not sure off the top of my head how I'd try to compute the overall availability #s from that one. One can possibly try to determine and sum the effects on the individual customers, but we can't from the information provided. But it's certainly less overall downtime than just counting it as a 7 hour failure.

Re: Post-Mortem for Google Compute Engine’s Global Outage on April 11

#55

> There are a number of lessons to be learned from this event -- for example, that the safeguard of a progressive rollout can be undone by a system designed to mask partial failures -- ... This is a really important point that should be more generally known. To quote Google's own "Paxos Made Live" paper, from 2007: > In closing we point out a challenge that we faced in testing our system for which we have no systemat…

This is an interesting question, and seems to get to the core of Nassim Taleb's ideas [1] about fragility and the limits of what we can understand, and how many of our attempts to create artificial stability ultimately bring about the opposite.

That said, based on this post-mortem, I think Google, and our industry as a whole, is doing a pretty good job. Periodic failures like this are inevitable, and if they serve to make it less likely that a similar failure occurs in the future, then that is a system as a whole that could be described as "anti-fragile".

[1] At least my interpretation of them

Re: Post-Mortem for Google Compute Engine’s Global Outage on April 11

#56

> There are a number of lessons to be learned from this event -- for example, that the safeguard of a progressive rollout can be undone by a system designed to mask partial failures -- ... This is a really important point that should be more generally known. To quote Google's own "Paxos Made Live" paper, from 2007: > In closing we point out a challenge that we faced in testing our system for which we have no systemat…

> So, has anyone managed to make progress toward a "systematic solution" in the last 9 years?

That depends on how you define "solution". If development time isn't a concern, then formal verification is a pretty solid solution. AWS has used TLA+ on a subset of its systems. [0]

[0] https://en.wikipedia.org/wiki/TLA%2B

Re: Post-Mortem for Google Compute Engine’s Global Outage on April 11

#57
post #49

Earlier quoted context omitted.

I know this logically. But emotionally I know how many bugs I have written in my life. I know software devs are human.. aka I know how the sausage is made.

Yes, but have you been part of the development & testing effort for mission-critical software (e.g. a class 1 or 2 medical device?). It's not true in all cases, but for the most part the level of QA that goes into the devices before release is significantly higher than that of your average product. This is why regulation is required.

[deleted]

Re: Post-Mortem for Google Compute Engine’s Global Outage on April 11

#58

This is a very good Post-Mortem. As I assumed it was kind of a corner case bug meet corner case bug met corner case bug. This is also why I am of afraid of a self driving cars and other such life critical software. There are going to be weird edge cases, what prevents you from reaching them? Making software is hard....

My consolation in that fact is that weird edge cases happen with human driven cars as well. Someone has a seizure and crashes, or more commonly reaches for a cigarette, the radio, their phone. People hit ice or water and overcorrect their spin. People drive too fast. Etc etc etc. Not even all edge cases, many common modes of failure. I except self-driving cars that kill people will be a huge emotional issue for a lot of people in accepting them, but for me, i just want them to be safer than human drivers, which isn't THAT high of a bar to cross.

Re: Post-Mortem for Google Compute Engine’s Global Outage on April 11

#59

This is very interesting. From the little I understand (sorry for using AWS terms as I am more versed with AWS than GCE) this can happen to AWS as well right? even if your software is deployed to multiple AZs / multiple regions, if bad routing / network configuration makes it through the various protection mechanisms then basically no amount of redundancy can help if your service is part of the non functional IP bloc…

Of course things such as mismatched account balances (i.e. the account balance does not equal all credits - debits) or erroneous postings due to bugs happen in the banking IT world. It's just that they are not that visible because only the people affected and that are quick enough in checking their balances will learn about it and after a few hours or days, when they notice the mistake, they will just fix the entries. (And if you thought you were clever and transferred all the funny money away, they are going to sue you to get it back; see e.g. this Quora thread for bank errors and legality: https://www.quora.com/If-my-bank-mistakenly-deposits-1-000-0...)

EDIT: Also, to answer the question: I think distributed computing is hard. The bank will usually have all their account balances on one huge central mainframe in one location, so you do not need to rely on computers talking to each other. And also, a bank does not really need to publish credits and debits at the same time - they just have to make sure your account is debited at or before the other account is credited (in fact, with most money transfers between banks there will be days between these two). So they can just debit your account, check whether this has worked and then send the money on its journey afterwards and be done with it. If a bug happens and the money does not show up at the recipient, they will complain, the bank can look into it and fix it - no (or not much to the bank, anyways) harm done.

Re: Post-Mortem for Google Compute Engine’s Global Outage on April 11

#60
post #6
post #3

Earlier quoted context omitted.

I wouldn't worry so much. I'm sure self driving cars are going to save a lot more lives than they are going to end. Humans are terrible drivers, and the software will only get better.

Yeah, remember, auto-pilot in a plane needs to be 100% reliable, or everyone dies. A car needs to be, I dunno, 80%? Compared to a bad human driver, who still drives every damn day, a computer need only be about 60% reliable to be better. People suck at driving. Even a shitty self-driving car will save a ton of lives simply by obeying traffic laws.

> Yeah, remember, auto-pilot in a plane needs to be 100% reliable, or everyone dies.

First, they're not anywhere near 100% reliable. They can fail on their own, and they'll also intentionally shut themselves off if the instruments they rely on fail. https://en.wikipedia.org/wiki/Air_France_Flight_447

Second, an autopilot failure shouldn't lead to death if the pilots are competent and paying attention.

Post reply on HN