AWS Service Disruption Post Mortem
aws.amazon.com
AWS Service Disruption Post Mortem
1–10 of 108 posts
Re: AWS Service Disruption Post Mortem
#2Re: AWS Service Disruption Post Mortem
#3Re: AWS Service Disruption Post Mortem
#4Re: AWS Service Disruption Post Mortem
#5Re: AWS Service Disruption Post Mortem
#6This supports the theory that between 50%-80% of outages are caused by human error, regardless of the resilience of the underlying infrastructure.
Re: AWS Service Disruption Post Mortem
#7I doubt this is the last time we'll hear of a "re-mirroring storm" in an oversaturated cloud.
Re: AWS Service Disruption Post Mortem
#8An automatic 100% credit for 10 days usage, thats pretty good IMO
Really the only purpose of a SLA penalty is to incentivize the provider to keep the network reliable.
Re: AWS Service Disruption Post Mortem
#9>...one of the standard steps is to shift traffic off of one of the redundant routers in the primary EBS network to allow the upgrade to happen. The traffic shift was executed incorrectly... This supports the theory that between 50%-80% of outages are caused by human error, regardless of the resilience of the underlying infrastructure.
Re: AWS Service Disruption Post Mortem
#10>...one of the standard steps is to shift traffic off of one of the redundant routers in the primary EBS network to allow the upgrade to happen. The traffic shift was executed incorrectly... This supports the theory that between 50%-80% of outages are caused by human error, regardless of the resilience of the underlying infrastructure.
Which leaves a question: why not engineer around humans, such that they are never needed in the day-to-day running of the systems?