Live data from Hacker News

AWS us-east-1 outage

status.aws.amazon.com

121–130 of 1001 posts

Re: AWS us-east-1 outage

#122

Earlier quoted context omitted.

Is this true or a joke? This sort of policy is how you destroy trust.

From what I've heard it's mostly true. Not only the CEO but a few SVPs can approve it, but yes a human must approve the update and it must be a high level exec. Part of the reason is because their SLAs are based on that dashboard, and that dashboard going red has a financial cost to AWS, so like any financial cost, it needs approval.

Sure, but... that just raises more questions :)

Taken literally what you are saying is the service could be down and an executive could override that, preventing them for paying customers for a service outage, even if the service did have an outage and the customer could prove it (screenshots, metrics from other cloud providers, many different folks see it).

I'm sure there is some subtlety to this, but it does mean that large corps with influence should be talking to AWS to ensure that status information corresponds with actual service outages.

Re: AWS us-east-1 outage

#124
post #90

Earlier quoted context omitted.

Man, some conclusions are being _jumped_ to by this reply.

There is a very long history of US-east-1 being horrible. Just bad. We've told every client we can to get out of there. It's one of the oldest amazon regions, and I think too much old legacy and weird stuff happens there. Use US-west-2.

Or US-East-2.

Re: AWS us-east-1 outage

#125

Earlier quoted context omitted.

I don't see why they couldn't provide an error rate graph like Reddit[0] or simply make services yellow saying "increased error rate detected, investigating..." 0: https://www.redditstatus.com/#system-metrics

Because Amazon has $$$$$ in their SLOs, and it costs them through the nose every minute they're down in payments made to customers and fees refunded. I trust them and most companies not to be outright fraudulent (although I'm sure some are), but it's totally understandable they'd be reticent to push the "Downtime Alert/Cost Us a Ton of Money" button until they're sure something serious is happening.

It literally is fraudulent though.

I don't think a region being down is something that you can be unsure about.

Re: AWS us-east-1 outage

#127
post #66

Are the actual services down, or is it just the console and/or login page? For example, the sign-up page appears to be working: https://portal.aws.amazon.com/billing/signup#/start Are websites that run on AWS us-east up? Are the AWS CLIs working?

I'm getting blank pages from Amazon.com itself.

Re: AWS us-east-1 outage

#128
post #38
post #13

I love that every time this happens, 100% of the services on https://status.aws.amazon.com are green.

Those five 9s don't come easy. Sometimes you have to prop them up :)

I wonder how often outages really happen. The official page is nonsense, of course, and we only collectively notice when the outage is big enough that lots of us are affected. On AWS, I see about a 3:1 ratio of "bump in the night" outages (quickly resolved, little corroboration) to mega too-big-to-hide outages. Does that mirror others' experiences?

Re: AWS us-east-1 outage

#129
post #66

Are the actual services down, or is it just the console and/or login page? For example, the sign-up page appears to be working: https://portal.aws.amazon.com/billing/signup#/start Are websites that run on AWS us-east up? Are the AWS CLIs working?

I can't access anything related to Cloudfront, either through the CLI or console :

  $ aws cloudfront list-distributions

  An error occurred (HttpTimeoutException) when calling the ListDistributions operation: Could not resolve DNS within remaining TTL of 4999 ms
However I can still access the distribution fine

Re: AWS us-east-1 outage

#130

It's funny that the first place I go to learn about the outage is Hacker News and not https://status.aws.amazon.com/ (it's still reports everything to be "operating normally"...)

I usually go on Twitter first for outages.
Post reply on HN