Live data from Hacker News

Summary of the Amazon S3 Service Disruption

aws.amazon.com

481–490 of 535 posts

Re: Summary of the Amazon S3 Service Disruption

#482

Earlier quoted context omitted.

It is unreasonable for me to think that company owners should have the spine to say, "We take the decision to fire someone very seriously. We'll take your comments under consideration, but we retain sole discretion over such decisions." It irks me that businesses fire people because of pressure from clients or social media. But having never been the boss, I may be missing something.

One reason to like a facet of Japanese management culture: if a customer wants someone to rake over the coals you offer management, not employees. Internal repercussions notwithstanding, externally the company is a united front. It cannot cause mistakes by luck, accident, or happenstance, because the world includes luck, accidents, and happenstance, so any user-visible error is ipso facto a failure of management.

The BBC staff have a term for the way the corporation almost does this: "deputy heads will roll"

Re: Summary of the Amazon S3 Service Disruption

#483
post #190

Imagine being THAT guy.......... in that exact moment...... after hitting enter and realizing what he did. RIP

https://blog.fastmail.com/2011/05/15/outage-report-a-cascade...

I was the guy who deployed the update to every one of our servers as I walked out the door for the day. So I know what this feels like.

You learn to get through it fast, because there's no other choice, you're almost always the best placed person to clean up the mess.

And one day you can look back and the scars have healed.

Re: Summary of the Amazon S3 Service Disruption

#484

Earlier quoted context omitted.

Any time I see "we're going to train everyone better" or "we're going to fire the guy who did it", all I can read is "this will happen again". You can't actually solve the problem of user error with training, and it's good to see Amazon not playing that game.

What bothered me about running TrueTech is that customers would sometimes demand repercussions against employees for making mistakes. Enter Frans Plugge. Whenever a customer would get into that mode we'd fire Frans. This was easy, simply because he didn't exist in the first place (his name was pulled from a skit by two Dutch comedians, bonus points if you know who and which skit). This usually then caused the custome…

Koot en Bie - Mannen Bellen

https://open.spotify.com/track/0tzlDFcFSNRUaaPnINWo8B

Re: Summary of the Amazon S3 Service Disruption

#485
post #400

Earlier quoted context omitted.

Agreed, especially regarding the culture but isn't this pretty much the same explanation they gave a few years ago when something similar happened? I seem to recall an EC2 or S3 outage a few years ago that boiled down to an engineer pushing out a patch that broke an entire region when it was supposed to be a phased deployment. I could be mis-remembering that but it's important that these lessons be applied across the…

Pretty sure that one was a Microsoft Azure outage. (Source: am a self-identified post-mortems connoisseur. :)

Do you by chance keep a public log of your postmortem collection :)?

Re: Summary of the Amazon S3 Service Disruption

#486

Earlier quoted context omitted.

A well-designed confirmation doesn't give you the same prompt for deleting some random test server as it does for deleting a production server. That helps with the "autopilot mode" issue.

I agree that it should help reduce the amount of mistakes. But I still believe auto-pilot mode is a real thing (and a danger!) . My point is that I'm not sure if it's even possible to design one that actually cuts errors to 0. And if that's indeed the case, even if it's close to 0, it's still non-zero, thus at the scale Amazon operates at, it's very probable that it will happen at least one time. Maybe sometime in th…

> My point is that I'm not sure if it's even possible to design one that actually cuts errors to 0.

Personally, I don't believe it is without making the tool impotent. But you can try and push down the error probability down to arbitrarily low value.

Re: Summary of the Amazon S3 Service Disruption

#487

Earlier quoted context omitted.

Any time I see "we're going to train everyone better" or "we're going to fire the guy who did it", all I can read is "this will happen again". You can't actually solve the problem of user error with training, and it's good to see Amazon not playing that game.

What bothered me about running TrueTech is that customers would sometimes demand repercussions against employees for making mistakes. Enter Frans Plugge. Whenever a customer would get into that mode we'd fire Frans. This was easy, simply because he didn't exist in the first place (his name was pulled from a skit by two Dutch comedians, bonus points if you know who and which skit). This usually then caused the custome…

Did you ever insist that Frans must be fired and refuse to accept the "we didn't mean it?"

Cause that sounds pretty great.

Re: Summary of the Amazon S3 Service Disruption

#488

Earlier quoted context omitted.

It's one of the reasons that silly guarantees like "twelve 9s of reliability" are meaningless. There are humans here. "Accidental human mishap" is gonna happen sometimes, and when it happens it's probably gonna affect a lot of data. Heck, at around 7 or 8 nines you have to account for the possibility that your operations team will decide that all your data is a vicious pack of timberwolves and needs to be defeated.

Note: they say "S3 is DESIGNED for 11 9s of durability". It's PR-speak to say that they don't give you any guarantee, but in theory the system is designed in a magnificent way.

Durability and uptime are not the same thing. Durability is about the chance of losing your data and has nothing to do with service disruptions. Their uptime SLA is much lower. Looking at [1], it looks like the SLA says 3 9s (discounts given for anything lower) of uptime.

[1] https://aws.amazon.com/s3/sla/

Re: Summary of the Amazon S3 Service Disruption

#489
post #176

Earlier quoted context omitted.

As I understand it, those guarantees don't mean that the service will actually stay up for the given number of 9s; it's that you'll be reimbursed monetarily if and when they go down.

I don't think it even means that; their policy says that the reimbursement only happens when your reliability dips all the way down to three 9s: https://aws.amazon.com/s3/sla/

The SLA (as you linked) says three nines. The 12 nines quoted by others is durability, not uptime.

Re: Summary of the Amazon S3 Service Disruption

#490
post #301

Earlier quoted context omitted.

This is true in some cases, but not when mitigations aren't practiced properly - it's not the fat fingered user who should be fired or retrained, but the designer or maintainer of the system that allowed it to become a serious issue. Look at the recent GitLab incident - one guy messed up and nuked a server. Okay, that happens sometimes, go to backups. Uh oh, all the backups are broken. Minor momentary problem just tu…

I don't buy the "If we just plan ENOUGH, disasters will never occur" argument. It's the universe is just too darn interesting for us to be able to plan enough to prevent it from being interesting.

"[T]he universe is just too darn interesting for us to be able to plan enough to prevent it from being interesting." -- Beat

That's a great line. How should I attribute it?

Post reply on HN