Live data from Hacker News

Post-Mortem for Google Compute Engine’s Global Outage on April 11

status.cloud.google.com

231–240 of 368 posts

Re: Post-Mortem for Google Compute Engine’s Global Outage on April 11

#231

Earlier quoted context omitted.

Self driving cars don't have to be perfect. They just have to be safer then driving is today [1]. The real question is if society can handle the unfairness that is death by random software error vs. death by negligent driving. It's easy to blame negligent driving on the driver, we're clearly not negligent so it really doesn't effect us right? But a software error might as well be an act of god, it's something that mi…

Well No, There is an upper limit on the damage a bad driver can do by say crushing his car with a bus or something like that. Imagine a bug or malware triggered at the same moment world-wide. It could kill millions. So it not as simple as 'It just has to be better than a human'

There is something that's called fail safe(ly). In case of any error inside the car, it should sowly decelerate and pull over. Sure, it will cause a lot of traffic problems if say 10% of all cars did that at the same time but the damage would not be as severe as a ghost driver entering the freeway with 140mph.

There are a few instances where some bad tesla batteries (the standard 12 volt batteries ironically) failed, and the cars handled it perfectly. It slowed down so that the driver could do safely pull over. Sure it did not happen to all cars at once, and autonomous cars migh not be able to do that by themselves but we have a log way to go to reach 100% autonomous driving (i.e. Without a steering wheel and a car that drives everywhere humans drive and not only on San Francisco's perfect sunny roads where it's been thoroughly tested on).

Re: Post-Mortem for Google Compute Engine’s Global Outage on April 11

#232
post #6

Earlier quoted context omitted.

Yeah, remember, auto-pilot in a plane needs to be 100% reliable, or everyone dies. A car needs to be, I dunno, 80%? Compared to a bad human driver, who still drives every damn day, a computer need only be about 60% reliable to be better. People suck at driving. Even a shitty self-driving car will save a ton of lives simply by obeying traffic laws.

> auto-pilot in a plane needs to be 100% reliable, or everyone dies Actually it is much more simpler than a self-driving car. And if there is a problem it disengages.

Autopilot in a plane usually doesn't involve autonavigation, whereas autodrive in a car generally requires navigation. Autodrive in a car without navigation is basically 'cruise control'.

Re: Post-Mortem for Google Compute Engine’s Global Outage on April 11

#233

Earlier quoted context omitted.

Well No, There is an upper limit on the damage a bad driver can do by say crushing his car with a bus or something like that. Imagine a bug or malware triggered at the same moment world-wide. It could kill millions. So it not as simple as 'It just has to be better than a human'

There is something that's called fail safe(ly). In case of any error inside the car, it should sowly decelerate and pull over. Sure, it will cause a lot of traffic problems if say 10% of all cars did that at the same time but the damage would not be as severe as a ghost driver entering the freeway with 140mph. There are a few instances where some bad tesla batteries (the standard 12 volt batteries ironically) failed,…

Google had fail safes which failed. What if a million cars pull over to the left instead of the right as a failed failsafe? Willing to risk your life?

Re: Post-Mortem for Google Compute Engine’s Global Outage on April 11

#234

Earlier quoted context omitted.

But if a new update changes how the car drives, wouldn't the data that would've been captured by the updated car be different than what the old outdated car recorded? For example if the new update makes the car more aggressive, then other real drivers might be more careful, slow down more, etc compared to the original runs?

I'd assume the more severe the change the more they'd want to test in the real world. And that once a comfortable self-driving experience is found, Google will not want to change it much.

> I'd assume the more severe the change the more they'd want to test in the real world.

I'm sorry but the important point is: how are we going to agree what needs more testing, and what can be updated without testing? If we let those big companies decide about those issues, then I'm afraid we will soon see another scandal like VW, except possibly with deadly consequences.

I can already predict the reasoning of those companies: last quarter our cars were safer than average, so we can afford some failures now.

Re: Post-Mortem for Google Compute Engine’s Global Outage on April 11

#235
post #201

Earlier quoted context omitted.

Self driving cars don't have to be perfect. They just have to be safer then driving is today [1]. The real question is if society can handle the unfairness that is death by random software error vs. death by negligent driving. It's easy to blame negligent driving on the driver, we're clearly not negligent so it really doesn't effect us right? But a software error might as well be an act of god, it's something that mi…

Well, this bug took down the entire system. What happens when self-driving software hits a similar bug? I don't think that there is any precedent for that sort of thing with manually driven cars. The scale could easily be larger than 100-car pile-ups due to poor weather conditions.

>I don't think that there is any precedent for that sort of thing with manually driven cars.

A stroke or heart attack while driving?

Re: Post-Mortem for Google Compute Engine’s Global Outage on April 11

#236

This is a very good Post-Mortem. As I assumed it was kind of a corner case bug meet corner case bug met corner case bug. This is also why I am of afraid of a self driving cars and other such life critical software. There are going to be weird edge cases, what prevents you from reaching them? Making software is hard....

Self driving cars don't have to be perfect. They just have to be safer then driving is today [1]. The real question is if society can handle the unfairness that is death by random software error vs. death by negligent driving. It's easy to blame negligent driving on the driver, we're clearly not negligent so it really doesn't effect us right? But a software error might as well be an act of god, it's something that mi…

There's also the ethical problem of life and death choices that will have to programmed in advance.

When an accident is inevitable, software will decide if prived or public property should be prioritized, which action is more likely to to protect driver/passenger A in detriment of driver/passenger B, etc.

Most people wouldn't blame the outcome of a split second decision made in heat of the moment but would take issue when the action is deliberate.

Interesting times we live in

Re: Post-Mortem for Google Compute Engine’s Global Outage on April 11

#237
post #214

Earlier quoted context omitted.

There are businesses that fit somewhere between Boeing and Spotify where failures still have some kind of steeper than casual cost. On Hacker News the "move fast and break things" ethos is probably making sense for many of the people submitting and commenting, since their business is closer to casual usage anyway. But that's not the whole audience.

Shit happens, when it comes to engineering, I'd trust Google more than even likely Boeing to manage systemic risk. As for cars, it's a real risk, but not the same as the bugs Google experienced; I personally have experienced a "bug" driving a car at high speeds, which resulted in a number of major electronic systems failing due to custom systems installed by a well known US startup.

Depends. You design and operate for a certain "shit happens" probability and price.

That's why I brought up airliners. You can't set low reliability goals and just say "shit happens". You would have less buyers, and it's not even legal anymore. So the bullet was bit and more reliable aircraft were developed. In the software world we're more like nineteen twenties still. That changed.

https://en.wikipedia.org/wiki/TWA_Flight_599

Let me phrase it in a perhaps less confrontational way. I see that there could be some business value in more reliable cloud platforms. There might be some business value with more nines in the availability percent, that is, less downtime per year. Or maybe just less global outages, even if that means more cases where some of a certain customers' containers or vm:s or what you have might be unavailable some of the time. That can be handled by running multiple units in the same cloud and with other techniques.

But at the moment, since there seem to be single points of failure (or policies that are single points of failure, like to update everything at once), if you, as a customer, would like to have more safety, you would have to run services in two different providers' cloud platforms. That could get slightly more complicated - and expensive as well. I guess some parts of these technologies are quite new so someone will come up with easy and good solutions.

Re: Post-Mortem for Google Compute Engine’s Global Outage on April 11

#238
post #20

Earlier quoted context omitted.

As long as the edge case bugs in self driving cars come up less frequently than human error, it's an overall improvement.

Currently, they don't. Google cars fail every 1,500 miles on average.

do you have a comparison for the number of miles between human errors?

Re: Post-Mortem for Google Compute Engine’s Global Outage on April 11

#239

Earlier quoted context omitted.

Currently, the self-driving software fails out on a Google Self-Driving Car every 1,500 miles. If the car suddenly stops trying to drive in the road, and the driver isn't attentive (or worse, if Google gets their way and convinces the laws to change so they don't have to have steering wheels) that's a lot of deaths. I'm not saying it won't get better, but pretending self-driving cars is a cure-all right now is hilari…

Source? What kind of failure are you talking about? Minor hiccups or full failures which stop the car entirely?

Google's report from December 2015: http://static.googleusercontent.com/media/www.google.com/en/...

Over 424,000 miles driven:

272 times the car had a 'system failure' and immediately returned control to the driver with only a couple seconds of warning. (Approx. every 1,558 miles.) A car mid-traffic spontaneously dropping control of the vehicle would likely create a large number of accidents.

13 car accidents prevented via human intervention (Approx. every 32,615 miles), 10 of which would've been the self-driving car's at-fault (Approx. every 42,400 miles). These virtual accidents were tested with the telemetry recorded during the incident, and it was determined had the human test driver not intervened, an accident would've occurred.

Total of these events is 285, which is approximately every 1,487 miles driven.

For useful comparison, a rough human average (when you add a large margin to account for unreported accidents) is somewhere around one accident every 150,000 miles driven. (Insurance companies see them every 250,000 miles approximately, I believe.)

Re: Post-Mortem for Google Compute Engine’s Global Outage on April 11

#240
post #238

Earlier quoted context omitted.

Currently, they don't. Google cars fail every 1,500 miles on average.

do you have a comparison for the number of miles between human errors?

It's hard to get an exact figure, particularly because of unreported accidents, and various sources. But I believe insurance companies have previously stated it's about one in every 250,000 miles. For the sake of giving a wide berth for unreported accidents, and to not give humans the benefit of the doubt, I've been using the rough figure of 150,000 miles between accidents.

I don't have a great source for it though, and if anyone finds a good source, it'd be fantastic.

Post reply on HN