Live data from Hacker News

Summary of the Amazon S3 Service Disruption

aws.amazon.com

521–530 of 535 posts

Re: Summary of the Amazon S3 Service Disruption

#521
post #409

Earlier quoted context omitted.

root@baz # shutdown now W: molly-guard: SSH session detected! Please type in hostname of the machine to shutdown: foo Good thing I asked; I won't shutdown baz ... Surprising to see such a simple protection neglected.

I don't know how well-known molly-guard is, but I've never heard of it before. Definitely enabling it on my servers next week.

Interesting. Up until now I've considered it well-known to the point of ubiquity :)

Re: Summary of the Amazon S3 Service Disruption

#522

Earlier quoted context omitted.

The BBC staff have a term for the way the corporation almost does this: "deputy heads will roll"

I always thought that was more a cynical take on the fact that the top guy was protected, rather than underlings.

Yes, exactly. Hence almost.

Re: Summary of the Amazon S3 Service Disruption

#524
post #354

Earlier quoted context omitted.

It occurs to me that having to type the English version of the numbers would probably work in this scenario. s3-shutdown -c "one hundred fifty" But something simpler like a --emergency flag or the more whimsical --shutitdownshutitalldown

Facebook uses a `--clowntown` flag to let people do things without normal checks. [0] [0] https://www.quora.com/When-people-who-work-at-Facebook-say-c...

I also rather like their intentionally verbose and scary parameter for generating arbitrary HTML in React: https://facebook.github.io/react/docs/dom-elements.html#dang...

Re: Summary of the Amazon S3 Service Disruption

#525
post #294

Earlier quoted context omitted.

I agree with that in general but having your monitoring system be dependent on the thing it monitors is a pretty big goof. It possible that the dependency was very non-obvious and many layers deep, which is more understandable, but still...its pretty fundamental.

The monitoring system was not dependent on the thing it was monitoring. The website that shows the public results of the monitoring, which is updates only by humans, depended on it.

I'm, I don't understand.

My us-east-1 RSS feed said S3 had no incidents.

Re: Summary of the Amazon S3 Service Disruption

#526
post #400
post #153

Earlier quoted context omitted.

Look at the language used though. This is saying very loudly "Look, this isn't the engineer's fault here". It's one thing I miss about Amazon's culture- not blaming people when system's fail. The follow-up doesn't bullshit with "extra training to make sure no one does this again", it says (effectively) "we're going to make this impossible to happen again, even if someone makes a mistake".

Agreed, especially regarding the culture but isn't this pretty much the same explanation they gave a few years ago when something similar happened? I seem to recall an EC2 or S3 outage a few years ago that boiled down to an engineer pushing out a patch that broke an entire region when it was supposed to be a phased deployment. I could be mis-remembering that but it's important that these lessons be applied across the…

Yeah an EC2 engineer switched over traffic to a backup network connection that had significantly less bandwidth, triggering cascading failures.

Re: Summary of the Amazon S3 Service Disruption

#527
post #450

Earlier quoted context omitted.

apparently this is (or was) a job in japan. companies would hire what amounts to an actor to get screamed at by the angry customer, and pretend to get fired on the spot. rinse, repeat whenever such appeasement is required.

I know one person who does this for real estate developers. He gets involved in contentious projects early on, goes to community meetings, offers testimony before the city council, etc. When construction gets going and people inevitably get pissed about some aspect of the project, he gets publicly fired to deflect the blame while the project moves on. Have seen it happen on three different projects in two cities now…

I don't know how to describe this in a single word or phrase appropriately, but I think it is a "genius problem" to exist. Not a genius solution. I feel that the problem itself is impressive and rich in layers of human nature, local culture etc - but once you have such a problem any average person could come up with a similar solution, because it is obvious.

It's still mind blowing and very amusing that this is a thing in our world!

Re: Summary of the Amazon S3 Service Disruption

#529
post #153

Earlier quoted context omitted.

Look at the language used though. This is saying very loudly "Look, this isn't the engineer's fault here". It's one thing I miss about Amazon's culture- not blaming people when system's fail. The follow-up doesn't bullshit with "extra training to make sure no one does this again", it says (effectively) "we're going to make this impossible to happen again, even if someone makes a mistake".

Any time I see "we're going to train everyone better" or "we're going to fire the guy who did it", all I can read is "this will happen again". You can't actually solve the problem of user error with training, and it's good to see Amazon not playing that game.

> You can't actually solve the problem of user error with training, and it's good to see Amazon not playing that game.

The problem of user error can be mitigated by an appropriate level of OCD.

But OCD can't be trained, you either have it or you don't.

Re: Summary of the Amazon S3 Service Disruption

#530
post #8

> At 9:37AM PST, an authorized S3 team member using an established playbook executed a command which was intended to remove a small number of servers for one of the S3 subsystems that is used by the S3 billing process. Unfortunately, one of the inputs to the command was entered incorrectly and a larger set of servers was removed than intended. It remains amazing to me that even with all the layers of automation, the…

serious question - why did no one ever accidentally launch and nuke a city, with thousands of nuclear warheads able to do so on short notice? like, AWS presumably puts a lot more redundancy in, and yet with all that effort comes up this far short. Why? It has a huge amount of brainpower all set up so that this never ever happens. Whatever works for the military, can't they adopt those actual best practices?

I highly recommend reading https://www.amazon.com/Command-Control-Damascus-Accident-Ill...

Turns out the answer to your question is simply: luck.

Post reply on HN