Live data from Hacker News

Summary of the Amazon S3 Service Disruption

aws.amazon.com

1–10 of 535 posts

Re: Summary of the Amazon S3 Service Disruption

#8
> At 9:37AM PST, an authorized S3 team member using an established playbook executed a command which was intended to remove a small number of servers for one of the S3 subsystems that is used by the S3 billing process. Unfortunately, one of the inputs to the command was entered incorrectly and a larger set of servers was removed than intended.

It remains amazing to me that even with all the layers of automation, the root cause of most serious deployment problems remain some variant of a fat fingered user.

Re: Summary of the Amazon S3 Service Disruption

#9
post #7
post #2

I wouldn't want to be the person who wrote the wrong command! Sheesh.

At these scales it's the fault of the system, not the individual, so hopefully they don't come down hard on them.

Agreed. They also seemed to acknowledge that in the post, as they mentioned improving the tool to not allow such destructive options.

Re: Summary of the Amazon S3 Service Disruption

#10
> From the beginning of this event until 11:37AM PST, we were unable to update the individual services’ status on the AWS Service Health Dashboard (SHD) because of a dependency the SHD administration console has on Amazon S3.

Ensuring that your status dashboard doesn't depend on the thing it's monitoring is probably the first thing you should think about when designing your status system. This doesn't fill me with confidence about how the rest of the system is designed, frankly...

Post reply on HN