Earlier quoted context omitted.
I'm pretty sure the sports meaning comes from the theathre meaning. Ansible has other bits of theatre metaphor in it, such as "roles" and "scripts".
What is the theatre meaning? I'm not familiar with that term.
Summary of the Amazon S3 Service Disruption
511–520 of 535 posts
Re: Summary of the Amazon S3 Service Disruption
#512Earlier quoted context omitted.
Imagine if the customer saw the same actor getting fired in different companies! Is the customer going to catch on? More likely, they will think "Yeah, no wonder there was a problem. This same incompetent dude wormed his way into this company too" :-)
There's a movie sketch in here somewhere. A guy has the worst day of his life, every single thing goes wrong, and at every single company the same person is "responsible" for the issue.
Re: Summary of the Amazon S3 Service Disruption
#513Earlier quoted context omitted.
What bothered me about running TrueTech is that customers would sometimes demand repercussions against employees for making mistakes. Enter Frans Plugge. Whenever a customer would get into that mode we'd fire Frans. This was easy, simply because he didn't exist in the first place (his name was pulled from a skit by two Dutch comedians, bonus points if you know who and which skit). This usually then caused the custome…
It is unreasonable for me to think that company owners should have the spine to say, "We take the decision to fire someone very seriously. We'll take your comments under consideration, but we retain sole discretion over such decisions." It irks me that businesses fire people because of pressure from clients or social media. But having never been the boss, I may be missing something.
Re: Summary of the Amazon S3 Service Disruption
#514Re: Summary of the Amazon S3 Service Disruption
#515Earlier quoted context omitted.
Pretty sure that one was a Microsoft Azure outage. (Source: am a self-identified post-mortems connoisseur. :)
Do you by chance keep a public log of your postmortem collection :)?
Re: Summary of the Amazon S3 Service Disruption
#516Earlier quoted context omitted.
That's what I meant in my original comment when I wrote that "number of hoops you have to jump through to do something should scale with the potential impact of an operation". Harmless operations - no confirmation. Something that could mess up your work - y-or-n-p confirmation. Something that could fuck up the whole infrastructure - you'd better get ready to retype a mix of "I DO UNDERSTAND WHAT I'M JUST ABOUT TO DO"…
Not sure if even that would work. I've almost deleted my heroku production server even though you need to type (or copy paste....ahem...) the full server name (e.g. thawing-temple-23345). I think the reason was that because in my mind I was 100% sure this was the right server, when the confirmation came up I didn't stopped to look if indeed this was the correct one so I mechanically started to type the name of the se…
Re: Summary of the Amazon S3 Service Disruption
#517> At 9:37AM PST, an authorized S3 team member using an established playbook executed a command which was intended to remove a small number of servers for one of the S3 subsystems that is used by the S3 billing process. Unfortunately, one of the inputs to the command was entered incorrectly and a larger set of servers was removed than intended. It remains amazing to me that even with all the layers of automation, the…
Re: Summary of the Amazon S3 Service Disruption
#518Earlier quoted context omitted.
Look at the language used though. This is saying very loudly "Look, this isn't the engineer's fault here". It's one thing I miss about Amazon's culture- not blaming people when system's fail. The follow-up doesn't bullshit with "extra training to make sure no one does this again", it says (effectively) "we're going to make this impossible to happen again, even if someone makes a mistake".
Any time I see "we're going to train everyone better" or "we're going to fire the guy who did it", all I can read is "this will happen again". You can't actually solve the problem of user error with training, and it's good to see Amazon not playing that game.
Re: Summary of the Amazon S3 Service Disruption
#519Earlier quoted context omitted.
To make error is human. To propagate error to all server in automatic way is #devops - DevOps Borat
I've long said something like "To err is human. To fuck up a million times in a second you need a computer." I may have to upgrade that to take the mighty power of Cloud (TM) into account, though. Billions and trillions of fuck ups per second are now well within reach!
Re: Summary of the Amazon S3 Service Disruption
#520" we have not completely restarted the index subsystem or the placement subsystem in our larger regions for many years. S3 has experienced massive growth over the last several years and the process of restarting these services and running the necessary safety checks to validate the integrity of the metadata took longer than expected" This is analogous to "we needed to fsck, and nobody realized how long that would tak…