Live data from Hacker News

AWS us-east-1 outage

status.aws.amazon.com

111–120 of 1001 posts

Re: AWS us-east-1 outage

#111

Earlier quoted context omitted.

EC2 or S3 showing red in any region literally requires personal approval of the CEO of AWS.

Is this true or a joke? This sort of policy is how you destroy trust.

From what I've heard it's mostly true. Not only the CEO but a few SVPs can approve it, but yes a human must approve the update and it must be a high level exec.

Part of the reason is because their SLAs are based on that dashboard, and that dashboard going red has a financial cost to AWS, so like any financial cost, it needs approval.

Re: AWS us-east-1 outage

#112
post #13

I love that every time this happens, 100% of the services on https://status.aws.amazon.com are green.

> Goodhart's Law is expressed simply as: “When a measure becomes a target, it ceases to be a good measure.”

It’s very frustrating. Why even have them?

Re: AWS us-east-1 outage

#115
post #13

I love that every time this happens, 100% of the services on https://status.aws.amazon.com are green.

On top of that, the "Personalized Health Dashboard" doesn't work because I can't seem to log in to the console.

I'm logged in; you're missing an error message.

Re: AWS us-east-1 outage

#116
post #13

I love that every time this happens, 100% of the services on https://status.aws.amazon.com are green.

EC2 or S3 showing red in any region literally requires personal approval of the CEO of AWS.

Unfortunately, errors don't require his approval...

Re: AWS us-east-1 outage

#117
I worked at a company that hired an ex-Amazon engineer to work on some cloud projects.

Whenever his projects went down, he fought tooth and nail against any suggestion to update the status page. When forced to update the status page, he'd follow up with an extremely long "post-mortem" document that was really just a long winded explanation about why the outage was someone else's fault.

He later explained that in his department at Amazon, being at fault for an outage was one of the worst things that could happen to you. He wanted to avoid that mark any way possible.

YMMV, of course. Amazon is a big company and I've had other friends work there in different departments who said this wasn't common at all. I will always remember the look of sheer panic he had when we insisted that he update the status page to accurately reflect an outage, though.

Re: AWS us-east-1 outage

#119
post #66

Are the actual services down, or is it just the console and/or login page? For example, the sign-up page appears to be working: https://portal.aws.amazon.com/billing/signup#/start Are websites that run on AWS us-east up? Are the AWS CLIs working?

Definitely not just the console. We had hundreds of thousands of websocket connections to us-east-1 drop at 15:40, and new websocket connections to that region are still failing. (Luckily not a huge impact on our service cause we run in 6 other regions, but still).

Re: AWS us-east-1 outage

#120

Earlier quoted context omitted.

Is this true or a joke? This sort of policy is how you destroy trust.

From what I've heard it's mostly true. Not only the CEO but a few SVPs can approve it, but yes a human must approve the update and it must be a high level exec. Part of the reason is because their SLAs are based on that dashboard, and that dashboard going red has a financial cost to AWS, so like any financial cost, it needs approval.

Being dishonest about SLAs seems to bear zero cost in this case?
Post reply on HN