>...one of the standard steps is to shift traffic off of one of the redundant routers in the primary EBS network to allow the upgrade to happen. The traffic shift was executed incorrectly... This supports the theory that between 50%-80% of outages are caused by human error, regardless of the resilience of the underlying infrastructure.
Which leaves a question: why not engineer around humans, such that they are never needed in the day-to-day running of the systems?
AWS Service Disruption Post Mortem
11–20 of 108 posts
Re: AWS Service Disruption Post Mortem
#12Earlier quoted context omitted.
Which leaves a question: why not engineer around humans, such that they are never needed in the day-to-day running of the systems?
"The trigger for this event was a network configuration change. We will audit our change process and increase the automation to prevent this mistake from happening in the future."
Re: AWS Service Disruption Post Mortem
#13>...one of the standard steps is to shift traffic off of one of the redundant routers in the primary EBS network to allow the upgrade to happen. The traffic shift was executed incorrectly... This supports the theory that between 50%-80% of outages are caused by human error, regardless of the resilience of the underlying infrastructure.
Which leaves a question: why not engineer around humans, such that they are never needed in the day-to-day running of the systems?
Re: AWS Service Disruption Post Mortem
#14Re: AWS Service Disruption Post Mortem
#15Lack in transparency in reaching out to customers is the biggest mistake what AWS did. They would learn from their mistakes, their servers and networks would be more reliable than ever.
This incident has given a reason for people to look at multi-cloud operation capability, for disaster recovery and backup reasons. AWS monopoly would be gone, there would be many new standards which would be proposed to bring in interoperability and for migrations between clouds.
Re: AWS Service Disruption Post Mortem
#16tl;dr version?
Re: AWS Service Disruption Post Mortem
#17AMZN has gotten a lot of flack over this outage, and rightly so. But I do want to dissuade anyone from thinking anybody else could do much better. I worked there 10 years ago, when they were closer to 200 engineers, and the caliber of people there at that point was insane. By far the smartest bunch I've ever worked with, and a place where I learned habits that serve me well to this day.
I know the guys that started the AWS group and they were the best of that already insanely selective group. It is easy to be an arm chair coach and scream that the network changes should have been automated in the first place, or that they should have predicted this storm, but that ignores just how fantastically hard what they are doing is and how fantastically well it works 99(how many 9's now?)% of the time.
In short, take my word for it, the people working on this are smarter than you and me, by an order of magnitude. There is no way you could do better, and it is unlikely that if you are building anything that needs more than a handful of servers you could build anything more reliable.
Re: AWS Service Disruption Post Mortem
#18During maintenance instead of shifting traffic off of one of the redundant routers the traffic was routed onto the lower capacity network. There was human error involved but the network issue only provoked latent bugs in the system that should have been picked out during disaster recovery testing.
Automatic recovery that isn't properly tested is a dangerous beast; it can cause problems faster and broader than any team of humans are capable of handling.
Re: AWS Service Disruption Post Mortem
#19tl;dr version?
Amazon offers a 10 day credit equal to 100% of their usage of EBS Volumes, EC2 Instances and RDS database instances. This credit will be automatically applied to the next bill.
Re: AWS Service Disruption Post Mortem
#20tldr: ""The trigger for this event was a network configuration change. We will audit our change process and increase the automation to prevent this mistake from happening in the future." AMZN has gotten a lot of flack over this outage, and rightly so. But I do want to dissuade anyone from thinking anybody else could do much better. I worked there 10 years ago, when they were closer to 200 engineers, and the caliber o…
If AWS with some of the smartest engineers can be down for that long do you think that our crappy service will be up 100% of time?