Live data from Hacker News

AWS us-east-1 outage

status.aws.amazon.com

151–160 of 1001 posts

Re: AWS us-east-1 outage

#151

Earlier quoted context omitted.

I don't see why they couldn't provide an error rate graph like Reddit[0] or simply make services yellow saying "increased error rate detected, investigating..." 0: https://www.redditstatus.com/#system-metrics

Because Amazon has $$$$$ in their SLOs, and it costs them through the nose every minute they're down in payments made to customers and fees refunded. I trust them and most companies not to be outright fraudulent (although I'm sure some are), but it's totally understandable they'd be reticent to push the "Downtime Alert/Cost Us a Ton of Money" button until they're sure something serious is happening.

It should be costing them trust not to push it when they should though. A trustworthy company will err on the side of pushing it. AWS is a near-monopoly, so their unprofessional business practices have still yet to cost them.

Re: AWS us-east-1 outage

#152

I worked at a company that hired an ex-Amazon engineer to work on some cloud projects. Whenever his projects went down, he fought tooth and nail against any suggestion to update the status page. When forced to update the status page, he'd follow up with an extremely long "post-mortem" document that was really just a long winded explanation about why the outage was someone else's fault. He later explained that in his…

On the retail/marketplace side this wasn't my experience, but we also didn't have any public dashboards. On Prime we occasionally had to refund in bulk, and when it was called for (internally or externally) we would right up a detailed post-mortem. This wasn't fun, but it was never about blaming a person and more about finding flaws in process or monitoring.

Re: AWS us-east-1 outage

#155
post #66

Are the actual services down, or is it just the console and/or login page? For example, the sign-up page appears to be working: https://portal.aws.amazon.com/billing/signup#/start Are websites that run on AWS us-east up? Are the AWS CLIs working?

I can tell you that some processes are not running, possibly due to SQS or SWF problems. Previous outages of this scale were caused by Kinesis outages. Can't connect via aws login at the cli either since we use SSO and that seems to be down.

Re: AWS us-east-1 outage

#156
Getting strange errors trying to manage my amazon account right now, could this be related?

494 ERROR and "We're sorry Something went wrong with our website, please try again later."

Re: AWS us-east-1 outage

#157
post #66

Are the actual services down, or is it just the console and/or login page? For example, the sign-up page appears to be working: https://portal.aws.amazon.com/billing/signup#/start Are websites that run on AWS us-east up? Are the AWS CLIs working?

EventBridge, CloudWatch. I've just started getting session errors with the console, too.

Re: AWS us-east-1 outage

#159

Earlier quoted context omitted.

On top of that, the "Personalized Health Dashboard" doesn't work because I can't seem to log in to the console.

I'm logged in; you're missing an error message.

We have federated login with MFA required (which was failing). It just started working again.

Scratch that... console is not loading at all now :)

Re: AWS us-east-1 outage

#160
post #65

We make heavy usage of Kinesis Firehose in us-east-1. Issues started ~1:24am ET and resolved around 7:31am ET. Then really kicked in at a much larger scale at 10:32am ET. We're now seeing failures with connections to RDS Postgres and other services. Console is completely unavailable to me.

Route53 is not updating new records. Console is also out.
Post reply on HN