Live data from Hacker News

AWS us-east-1 outage

status.aws.amazon.com

181–190 of 1001 posts

Re: AWS us-east-1 outage

#182

Earlier quoted context omitted.

Because Amazon has $$$$$ in their SLOs, and it costs them through the nose every minute they're down in payments made to customers and fees refunded. I trust them and most companies not to be outright fraudulent (although I'm sure some are), but it's totally understandable they'd be reticent to push the "Downtime Alert/Cost Us a Ton of Money" button until they're sure something serious is happening.

It literally is fraudulent though. I don't think a region being down is something that you can be unsure about.

Oh, you can get pretty weaselly about what “down” means. If there is “just” an S3 issue, are all the various services which are still “available” but throwing an elevated number of errors because of their own internal dependency on S3 actually down or just “degraded?” You have to spin up the hair-splitting apparatus early in the incident to try to keep clear of the post-mortem party. :D

Re: AWS us-east-1 outage

#183
post #13

I love that every time this happens, 100% of the services on https://status.aws.amazon.com are green.

Even better, when I try to go to console, I get:

> AWS Management Console Home page is currently unavailable.

> You can monitor status on the AWS Service Health Dashboard.

"AWS Service Health Dashboard" is a link to status.aws.amazon.com... which is ALL GREEN. So... thanks for the suggestion?

At this point the AWS service health dashboard is kind of famous for always been green isn't it? It's a joke to it's users. Do the folks who work on the relevant AWS internal team(s) know this, and just not have the resources to do anything about it, or what? If it's a harder problem than you'd think for interesting technical reasons, that'd be interesting to hear about.

Re: AWS us-east-1 outage

#184
post #62

Azure, Google Cloud, AWS and others need to have a “Status alliance” where they determine the status of each of their services by a quorum using all cloud providers. Status pages are virtually useless these days

They can do this without an alliance. They very intentionally choose not to do it. Every major company has moved away from having accurate status pages.

It's because none of these companies are held responsible for missing their actual SLAs, as opposed to their self-reported SLA compliance.

So unless regulation gets implemented that says otherwise, there's zero incentive for any company to maintain an accurate status page.

Re: AWS us-east-1 outage

#185

I worked at a company that hired an ex-Amazon engineer to work on some cloud projects. Whenever his projects went down, he fought tooth and nail against any suggestion to update the status page. When forced to update the status page, he'd follow up with an extremely long "post-mortem" document that was really just a long winded explanation about why the outage was someone else's fault. He later explained that in his…

This gets posted every time there's an AWS outage. It mind as well be a copy pasta at this point.

well, this is the first time I've seen it, so I am glad it was posted this time.

Re: AWS us-east-1 outage

#188
post #113

Friends tell friends to pick us-east-2. Virginia is for lovers, Ohio is for availability.

Lots of services are only in us-east-1. The sso system isn't working 100% right now so that's where I assume it's hosted.

Yeah I can't log in with our external SAML SSO to our AWS dashboard to manage our us-east-2 resources. . . . Because our auth is apparently routed thru us-east-1 STS

Re: AWS us-east-1 outage

#189
post #122

Earlier quoted context omitted.

From what I've heard it's mostly true. Not only the CEO but a few SVPs can approve it, but yes a human must approve the update and it must be a high level exec. Part of the reason is because their SLAs are based on that dashboard, and that dashboard going red has a financial cost to AWS, so like any financial cost, it needs approval.

Sure, but... that just raises more questions :) Taken literally what you are saying is the service could be down and an executive could override that, preventing them for paying customers for a service outage, even if the service did have an outage and the customer could prove it (screenshots, metrics from other cloud providers, many different folks see it). I'm sure there is some subtlety to this, but it does mean t…

Large corps with influence get what they want regardless. Status page goes red and the small corps start thinking they can get what they want too.

Re: AWS us-east-1 outage

#190

I worked at a company that hired an ex-Amazon engineer to work on some cloud projects. Whenever his projects went down, he fought tooth and nail against any suggestion to update the status page. When forced to update the status page, he'd follow up with an extremely long "post-mortem" document that was really just a long winded explanation about why the outage was someone else's fault. He later explained that in his…

This gets posted every time there's an AWS outage. It mind as well be a copy pasta at this point.

[deleted]
Post reply on HN