Summary of the Amazon S3 Service Disruption
481–490 of 535 posts
Re: Summary of the Amazon S3 Service Disruption
#482Earlier quoted context omitted.
It is unreasonable for me to think that company owners should have the spine to say, "We take the decision to fire someone very seriously. We'll take your comments under consideration, but we retain sole discretion over such decisions." It irks me that businesses fire people because of pressure from clients or social media. But having never been the boss, I may be missing something.
One reason to like a facet of Japanese management culture: if a customer wants someone to rake over the coals you offer management, not employees. Internal repercussions notwithstanding, externally the company is a united front. It cannot cause mistakes by luck, accident, or happenstance, because the world includes luck, accidents, and happenstance, so any user-visible error is ipso facto a failure of management.
Re: Summary of the Amazon S3 Service Disruption
#483Imagine being THAT guy.......... in that exact moment...... after hitting enter and realizing what he did. RIP
I was the guy who deployed the update to every one of our servers as I walked out the door for the day. So I know what this feels like.
You learn to get through it fast, because there's no other choice, you're almost always the best placed person to clean up the mess.
And one day you can look back and the scars have healed.
Re: Summary of the Amazon S3 Service Disruption
#484Earlier quoted context omitted.
Any time I see "we're going to train everyone better" or "we're going to fire the guy who did it", all I can read is "this will happen again". You can't actually solve the problem of user error with training, and it's good to see Amazon not playing that game.
What bothered me about running TrueTech is that customers would sometimes demand repercussions against employees for making mistakes. Enter Frans Plugge. Whenever a customer would get into that mode we'd fire Frans. This was easy, simply because he didn't exist in the first place (his name was pulled from a skit by two Dutch comedians, bonus points if you know who and which skit). This usually then caused the custome…
Re: Summary of the Amazon S3 Service Disruption
#485Earlier quoted context omitted.
Agreed, especially regarding the culture but isn't this pretty much the same explanation they gave a few years ago when something similar happened? I seem to recall an EC2 or S3 outage a few years ago that boiled down to an engineer pushing out a patch that broke an entire region when it was supposed to be a phased deployment. I could be mis-remembering that but it's important that these lessons be applied across the…
Pretty sure that one was a Microsoft Azure outage. (Source: am a self-identified post-mortems connoisseur. :)
Re: Summary of the Amazon S3 Service Disruption
#486Earlier quoted context omitted.
A well-designed confirmation doesn't give you the same prompt for deleting some random test server as it does for deleting a production server. That helps with the "autopilot mode" issue.
I agree that it should help reduce the amount of mistakes. But I still believe auto-pilot mode is a real thing (and a danger!) . My point is that I'm not sure if it's even possible to design one that actually cuts errors to 0. And if that's indeed the case, even if it's close to 0, it's still non-zero, thus at the scale Amazon operates at, it's very probable that it will happen at least one time. Maybe sometime in th…
Personally, I don't believe it is without making the tool impotent. But you can try and push down the error probability down to arbitrarily low value.
Re: Summary of the Amazon S3 Service Disruption
#487Earlier quoted context omitted.
Any time I see "we're going to train everyone better" or "we're going to fire the guy who did it", all I can read is "this will happen again". You can't actually solve the problem of user error with training, and it's good to see Amazon not playing that game.
What bothered me about running TrueTech is that customers would sometimes demand repercussions against employees for making mistakes. Enter Frans Plugge. Whenever a customer would get into that mode we'd fire Frans. This was easy, simply because he didn't exist in the first place (his name was pulled from a skit by two Dutch comedians, bonus points if you know who and which skit). This usually then caused the custome…
Cause that sounds pretty great.
Re: Summary of the Amazon S3 Service Disruption
#488Earlier quoted context omitted.
It's one of the reasons that silly guarantees like "twelve 9s of reliability" are meaningless. There are humans here. "Accidental human mishap" is gonna happen sometimes, and when it happens it's probably gonna affect a lot of data. Heck, at around 7 or 8 nines you have to account for the possibility that your operations team will decide that all your data is a vicious pack of timberwolves and needs to be defeated.
Note: they say "S3 is DESIGNED for 11 9s of durability". It's PR-speak to say that they don't give you any guarantee, but in theory the system is designed in a magnificent way.
Re: Summary of the Amazon S3 Service Disruption
#489Earlier quoted context omitted.
As I understand it, those guarantees don't mean that the service will actually stay up for the given number of 9s; it's that you'll be reimbursed monetarily if and when they go down.
I don't think it even means that; their policy says that the reimbursement only happens when your reliability dips all the way down to three 9s: https://aws.amazon.com/s3/sla/
Re: Summary of the Amazon S3 Service Disruption
#490Earlier quoted context omitted.
This is true in some cases, but not when mitigations aren't practiced properly - it's not the fat fingered user who should be fired or retrained, but the designer or maintainer of the system that allowed it to become a serious issue. Look at the recent GitLab incident - one guy messed up and nuked a server. Okay, that happens sometimes, go to backups. Uh oh, all the backups are broken. Minor momentary problem just tu…
I don't buy the "If we just plan ENOUGH, disasters will never occur" argument. It's the universe is just too darn interesting for us to be able to plan enough to prevent it from being interesting.
That's a great line. How should I attribute it?