I love that every time this happens, 100% of the services on https://status.aws.amazon.com are green.
Status pages are hard
AWS us-east-1 outage
221–230 of 1001 posts
Re: AWS us-east-1 outage
#222Earlier quoted context omitted.
This gets posted every time there's an AWS outage. It mind as well be a copy pasta at this point.
well, this is the first time I've seen it, so I am glad it was posted this time.
Re: AWS us-east-1 outage
#223I'm now getting failures searching for products on Amazon.com itself. This is somewhat surprising, as the narrative always was that Amazon didn't do a great job of dogfooding their own cloud platform.
Re: AWS us-east-1 outage
#224Earlier quoted context omitted.
From what I've heard it's mostly true. Not only the CEO but a few SVPs can approve it, but yes a human must approve the update and it must be a high level exec. Part of the reason is because their SLAs are based on that dashboard, and that dashboard going red has a financial cost to AWS, so like any financial cost, it needs approval.
Sure, but... that just raises more questions :) Taken literally what you are saying is the service could be down and an executive could override that, preventing them for paying customers for a service outage, even if the service did have an outage and the customer could prove it (screenshots, metrics from other cloud providers, many different folks see it). I'm sure there is some subtlety to this, but it does mean t…
Re: AWS us-east-1 outage
#225Re: AWS us-east-1 outage
#226Re: AWS us-east-1 outage
#227I worked at a company that hired an ex-Amazon engineer to work on some cloud projects. Whenever his projects went down, he fought tooth and nail against any suggestion to update the status page. When forced to update the status page, he'd follow up with an extremely long "post-mortem" document that was really just a long winded explanation about why the outage was someone else's fault. He later explained that in his…
That sounds like the exact opposite of human-factors engineering. No one likes taking blame. But when things go sideways, people are extra spicy and defensive, which makes them clam up and often withhold useful information, which can extend the outage. No-blame analysis is a much better pattern. Everyone wins. It's about building the system that builds the system. Stuff broke; fix the stuff that broke, then fix the t…
Re: AWS us-east-1 outage
#228ECR borked for us in east-1
Re: AWS us-east-1 outage
#229Are the actual services down, or is it just the console and/or login page? For example, the sign-up page appears to be working: https://portal.aws.amazon.com/billing/signup#/start Are websites that run on AWS us-east up? Are the AWS CLIs working?
Definitely not just the console. We had hundreds of thousands of websocket connections to us-east-1 drop at 15:40, and new websocket connections to that region are still failing. (Luckily not a huge impact on our service cause we run in 6 other regions, but still).
Re: AWS us-east-1 outage
#230I love that every time this happens, 100% of the services on https://status.aws.amazon.com are green.
Those five 9s don't come easy. Sometimes you have to prop them up :)