Live data from Hacker News

Summary of the Amazon S3 Service Disruption

aws.amazon.com

41–50 of 535 posts

Re: Summary of the Amazon S3 Service Disruption

#41
For the many of us who have built businesses dependent on S3, is anyone else surprised at a few assumptions embedded here?

* "authorized S3 team member" -- how did this team member acquire these elevated privs?

* Running playbooks is done by one member without a second set of eyes or approval?

* "we have not completely restarted the index subsystem or the placement subsystem in our larger regions for many years"

The good news:

* "The S3 team had planned further partitioning of the index subsystem later this year. We are reprioritizing that work to begin immediately."

The truly embarrassing that everyone has known about for years is the status page:

* "we were unable to update the individual services’ status on the AWS Service Health Dashboard "

When there is a wildly-popular Chrome plugin to fix your page ("Real AWS Status") you would think a company as responsive as AWS would have fixed this years ago.

Re: Summary of the Amazon S3 Service Disruption

#43
post #25
post #8

> At 9:37AM PST, an authorized S3 team member using an established playbook executed a command which was intended to remove a small number of servers for one of the S3 subsystems that is used by the S3 billing process. Unfortunately, one of the inputs to the command was entered incorrectly and a larger set of servers was removed than intended. It remains amazing to me that even with all the layers of automation, the…

Wonder what happened to the poor slob who did that. He/She was unauthorized AND caused a pretty serious outage... EDIT: derp, my bad, I read as "unauthorized" which was "authorized".

I think the article says he was authorized? Either way, at worst, he got fired.

Re: Summary of the Amazon S3 Service Disruption

#44
post #11

What's missing is addressing the problems with their status page system, and how we all had to use Hacker News and other sources to confirm that US East was borked.

This was addressed in the post:

"From the beginning of this event until 11:37AM PST, we were unable to update the individual services’ status on the AWS Service Health Dashboard (SHD) because of a dependency the SHD administration console has on Amazon S3. Instead, we used the AWS Twitter feed (@AWSCloud) and SHD banner text to communicate status until we were able to update the individual services’ status on the SHD. We understand that the SHD provides important visibility to our customers during operational events and we have changed the SHD administration console to run across multiple AWS regions."

Re: Summary of the Amazon S3 Service Disruption

#45

For the many of us who have built businesses dependent on S3, is anyone else surprised at a few assumptions embedded here? * "authorized S3 team member" -- how did this team member acquire these elevated privs? * Running playbooks is done by one member without a second set of eyes or approval? * "we have not completely restarted the index subsystem or the placement subsystem in our larger regions for many years" The…

>Unauthorized S3 team member

This is not in the post.

Re: Summary of the Amazon S3 Service Disruption

#46
post #25
post #8

> At 9:37AM PST, an authorized S3 team member using an established playbook executed a command which was intended to remove a small number of servers for one of the S3 subsystems that is used by the S3 billing process. Unfortunately, one of the inputs to the command was entered incorrectly and a larger set of servers was removed than intended. It remains amazing to me that even with all the layers of automation, the…

Wonder what happened to the poor slob who did that. He/She was unauthorized AND caused a pretty serious outage... EDIT: derp, my bad, I read as "unauthorized" which was "authorized".

> At 9:37AM PST, an >>> authorized [...]

typos happen, if the system didn't stop them that's a design issue or accepted risk.

Re: Summary of the Amazon S3 Service Disruption

#47
post #25
post #8

> At 9:37AM PST, an authorized S3 team member using an established playbook executed a command which was intended to remove a small number of servers for one of the S3 subsystems that is used by the S3 billing process. Unfortunately, one of the inputs to the command was entered incorrectly and a larger set of servers was removed than intended. It remains amazing to me that even with all the layers of automation, the…

Wonder what happened to the poor slob who did that. He/She was unauthorized AND caused a pretty serious outage... EDIT: derp, my bad, I read as "unauthorized" which was "authorized".

Did the parent comment change? It currently says "an authorized..."

Re: Summary of the Amazon S3 Service Disruption

#48

For the many of us who have built businesses dependent on S3, is anyone else surprised at a few assumptions embedded here? * "authorized S3 team member" -- how did this team member acquire these elevated privs? * Running playbooks is done by one member without a second set of eyes or approval? * "we have not completely restarted the index subsystem or the placement subsystem in our larger regions for many years" The…

) An authorized, not unauthorized. They sound almost identical, English language is... Yea :/

) If it's a playbook for something with minimal intended impact sure. The issue is that the tooling had larger capabilities than should be.

) Yes that seems like a major, major problem.

Re: Summary of the Amazon S3 Service Disruption

#49

For the many of us who have built businesses dependent on S3, is anyone else surprised at a few assumptions embedded here? * "authorized S3 team member" -- how did this team member acquire these elevated privs? * Running playbooks is done by one member without a second set of eyes or approval? * "we have not completely restarted the index subsystem or the placement subsystem in our larger regions for many years" The…

[deleted]

Re: Summary of the Amazon S3 Service Disruption

#50

For the many of us who have built businesses dependent on S3, is anyone else surprised at a few assumptions embedded here? * "authorized S3 team member" -- how did this team member acquire these elevated privs? * Running playbooks is done by one member without a second set of eyes or approval? * "we have not completely restarted the index subsystem or the placement subsystem in our larger regions for many years" The…

It says authorized, not unauthorized.
Post reply on HN