Live data from Hacker News

Summary of the Amazon S3 Service Disruption

aws.amazon.com

511–520 of 535 posts

Re: Summary of the Amazon S3 Service Disruption

#511

Earlier quoted context omitted.

I'm pretty sure the sports meaning comes from the theathre meaning. Ansible has other bits of theatre metaphor in it, such as "roles" and "scripts".

What is the theatre meaning? I'm not familiar with that term.

It's just a book that contains one or more scripts, for example: https://en.wikipedia.org/wiki/Fleury_Playbook

Re: Summary of the Amazon S3 Service Disruption

#512

Earlier quoted context omitted.

Imagine if the customer saw the same actor getting fired in different companies! Is the customer going to catch on? More likely, they will think "Yeah, no wonder there was a problem. This same incompetent dude wormed his way into this company too" :-)

There's a movie sketch in here somewhere. A guy has the worst day of his life, every single thing goes wrong, and at every single company the same person is "responsible" for the issue.

Awesome! The first movie ever made based on an anonymous comment on HN. Wait. So I can't get a cut in the profits then?

Re: Summary of the Amazon S3 Service Disruption

#513

Earlier quoted context omitted.

What bothered me about running TrueTech is that customers would sometimes demand repercussions against employees for making mistakes. Enter Frans Plugge. Whenever a customer would get into that mode we'd fire Frans. This was easy, simply because he didn't exist in the first place (his name was pulled from a skit by two Dutch comedians, bonus points if you know who and which skit). This usually then caused the custome…

It is unreasonable for me to think that company owners should have the spine to say, "We take the decision to fire someone very seriously. We'll take your comments under consideration, but we retain sole discretion over such decisions." It irks me that businesses fire people because of pressure from clients or social media. But having never been the boss, I may be missing something.

On the contrary, I think it's quite reasonable.

Re: Summary of the Amazon S3 Service Disruption

#515
post #485

Earlier quoted context omitted.

Pretty sure that one was a Microsoft Azure outage. (Source: am a self-identified post-mortems connoisseur. :)

Do you by chance keep a public log of your postmortem collection :)?

I don't, but danluu does! https://github.com/danluu/post-mortems

Re: Summary of the Amazon S3 Service Disruption

#516

Earlier quoted context omitted.

That's what I meant in my original comment when I wrote that "number of hoops you have to jump through to do something should scale with the potential impact of an operation". Harmless operations - no confirmation. Something that could mess up your work - y-or-n-p confirmation. Something that could fuck up the whole infrastructure - you'd better get ready to retype a mix of "I DO UNDERSTAND WHAT I'M JUST ABOUT TO DO"…

Not sure if even that would work. I've almost deleted my heroku production server even though you need to type (or copy paste....ahem...) the full server name (e.g. thawing-temple-23345). I think the reason was that because in my mind I was 100% sure this was the right server, when the confirmation came up I didn't stopped to look if indeed this was the correct one so I mechanically started to type the name of the se…

Still no help against "whoops, took down a different production instance than intended."

Re: Summary of the Amazon S3 Service Disruption

#517
post #8

> At 9:37AM PST, an authorized S3 team member using an established playbook executed a command which was intended to remove a small number of servers for one of the S3 subsystems that is used by the S3 billing process. Unfortunately, one of the inputs to the command was entered incorrectly and a larger set of servers was removed than intended. It remains amazing to me that even with all the layers of automation, the…

20 years ago I read a postmortem of Tandem and their Non-Stop Unix. A core take-away for me was: "Computer hardware has gotten way more reliable than it was." combined with "The leading cause of outages has become operators making mistakes."

Re: Summary of the Amazon S3 Service Disruption

#518
post #153

Earlier quoted context omitted.

Look at the language used though. This is saying very loudly "Look, this isn't the engineer's fault here". It's one thing I miss about Amazon's culture- not blaming people when system's fail. The follow-up doesn't bullshit with "extra training to make sure no one does this again", it says (effectively) "we're going to make this impossible to happen again, even if someone makes a mistake".

Any time I see "we're going to train everyone better" or "we're going to fire the guy who did it", all I can read is "this will happen again". You can't actually solve the problem of user error with training, and it's good to see Amazon not playing that game.

Years ago I read a story about a fat-fingered ops person getting called into the CEO's office after an outage. "I thought you were calling me in to fire me." "I can't afford to fire you, today I spent a million dollars training you."

Re: Summary of the Amazon S3 Service Disruption

#519
post #73
post #65

Earlier quoted context omitted.

To make error is human. To propagate error to all server in automatic way is #devops - DevOps Borat

I've long said something like "To err is human. To fuck up a million times in a second you need a computer." I may have to upgrade that to take the mighty power of Cloud (TM) into account, though. Billions and trillions of fuck ups per second are now well within reach!

"A computer lets you make more mistakes faster than any other invention with the possible exception of handguns and tequila." -- Mitch Ratcliffe

Re: Summary of the Amazon S3 Service Disruption

#520
post #40

" we have not completely restarted the index subsystem or the placement subsystem in our larger regions for many years. S3 has experienced massive growth over the last several years and the process of restarting these services and running the necessary safety checks to validate the integrity of the metadata took longer than expected" This is analogous to "we needed to fsck, and nobody realized how long that would tak…

Reminds me of the last time I borked my btrfs.
Post reply on HN