Earlier quoted context omitted.
root@baz # shutdown now W: molly-guard: SSH session detected! Please type in hostname of the machine to shutdown: foo Good thing I asked; I won't shutdown baz ... Surprising to see such a simple protection neglected.
I don't know how well-known molly-guard is, but I've never heard of it before. Definitely enabling it on my servers next week.
Summary of the Amazon S3 Service Disruption
521–530 of 535 posts
Re: Summary of the Amazon S3 Service Disruption
#522Re: Summary of the Amazon S3 Service Disruption
#523Re: Summary of the Amazon S3 Service Disruption
#524Earlier quoted context omitted.
It occurs to me that having to type the English version of the numbers would probably work in this scenario. s3-shutdown -c "one hundred fifty" But something simpler like a --emergency flag or the more whimsical --shutitdownshutitalldown
Facebook uses a `--clowntown` flag to let people do things without normal checks. [0] [0] https://www.quora.com/When-people-who-work-at-Facebook-say-c...
Re: Summary of the Amazon S3 Service Disruption
#525Earlier quoted context omitted.
I agree with that in general but having your monitoring system be dependent on the thing it monitors is a pretty big goof. It possible that the dependency was very non-obvious and many layers deep, which is more understandable, but still...its pretty fundamental.
The monitoring system was not dependent on the thing it was monitoring. The website that shows the public results of the monitoring, which is updates only by humans, depended on it.
My us-east-1 RSS feed said S3 had no incidents.
Re: Summary of the Amazon S3 Service Disruption
#526Earlier quoted context omitted.
Look at the language used though. This is saying very loudly "Look, this isn't the engineer's fault here". It's one thing I miss about Amazon's culture- not blaming people when system's fail. The follow-up doesn't bullshit with "extra training to make sure no one does this again", it says (effectively) "we're going to make this impossible to happen again, even if someone makes a mistake".
Agreed, especially regarding the culture but isn't this pretty much the same explanation they gave a few years ago when something similar happened? I seem to recall an EC2 or S3 outage a few years ago that boiled down to an engineer pushing out a patch that broke an entire region when it was supposed to be a phased deployment. I could be mis-remembering that but it's important that these lessons be applied across the…
Re: Summary of the Amazon S3 Service Disruption
#527Earlier quoted context omitted.
apparently this is (or was) a job in japan. companies would hire what amounts to an actor to get screamed at by the angry customer, and pretend to get fired on the spot. rinse, repeat whenever such appeasement is required.
I know one person who does this for real estate developers. He gets involved in contentious projects early on, goes to community meetings, offers testimony before the city council, etc. When construction gets going and people inevitably get pissed about some aspect of the project, he gets publicly fired to deflect the blame while the project moves on. Have seen it happen on three different projects in two cities now…
It's still mind blowing and very amusing that this is a thing in our world!
Re: Summary of the Amazon S3 Service Disruption
#528Re: Summary of the Amazon S3 Service Disruption
#529Earlier quoted context omitted.
Look at the language used though. This is saying very loudly "Look, this isn't the engineer's fault here". It's one thing I miss about Amazon's culture- not blaming people when system's fail. The follow-up doesn't bullshit with "extra training to make sure no one does this again", it says (effectively) "we're going to make this impossible to happen again, even if someone makes a mistake".
Any time I see "we're going to train everyone better" or "we're going to fire the guy who did it", all I can read is "this will happen again". You can't actually solve the problem of user error with training, and it's good to see Amazon not playing that game.
The problem of user error can be mitigated by an appropriate level of OCD.
But OCD can't be trained, you either have it or you don't.
Re: Summary of the Amazon S3 Service Disruption
#530> At 9:37AM PST, an authorized S3 team member using an established playbook executed a command which was intended to remove a small number of servers for one of the S3 subsystems that is used by the S3 billing process. Unfortunately, one of the inputs to the command was entered incorrectly and a larger set of servers was removed than intended. It remains amazing to me that even with all the layers of automation, the…
serious question - why did no one ever accidentally launch and nuke a city, with thousands of nuclear warheads able to do so on short notice? like, AWS presumably puts a lot more redundancy in, and yet with all that effort comes up this far short. Why? It has a huge amount of brainpower all set up so that this never ever happens. Whatever works for the military, can't they adopt those actual best practices?
Turns out the answer to your question is simply: luck.