I love that every time this happens, 100% of the services on https://status.aws.amazon.com are green.
I don't see why they couldn't provide an error rate graph like Reddit[0] or simply make services yellow saying "increased error rate detected, investigating..." 0: https://www.redditstatus.com/#system-metrics
AWS us-east-1 outage
511–520 of 1001 posts
Re: AWS us-east-1 outage
#512Re: AWS us-east-1 outage
#513My job (although 50% of time) at Azure is unit testing/monitoring services under different scenarios and flows to detect small failures that will be overlooked in public status page. Our tests run multiple times daily and we have people constantly monitoring logs. It concerns me when I see all AWS services are 100% green when I know there is an outage.
Re: AWS us-east-1 outage
#514I worked at a company that hired an ex-Amazon engineer to work on some cloud projects. Whenever his projects went down, he fought tooth and nail against any suggestion to update the status page. When forced to update the status page, he'd follow up with an extremely long "post-mortem" document that was really just a long winded explanation about why the outage was someone else's fault. He later explained that in his…
It's popular to upvote this during outages, because it fits a narrative. The truth (as always) is more complex: * No, this isn't the broad culture. It's not even a blip. These are EXCEPTIONAL circumstances by extremely bad teams that - if and when found out - would be intervened dramatically. * The broad culture is blameless post-mortems. Not whose fault is it. But what was the problem and how to fix it. And one of t…
See, AWS is basically turning into a long standing utility that needs to be reliable.
Hey, do most institutions like that completely turn over their staff every three years? Yeah, no.
Great for building it out and grabbing market share.
Maybe not for being the basis of a reliable substrate of the modern internet.
If there are dozens of bespoke systems that keep AWS afloat (disclosure: I have friends who worked there, and there are, and also Conway's law), but if the people who wrote them are three generations of HIRE/FIRE ago....
Not good.
Re: AWS us-east-1 outage
#515Re: AWS us-east-1 outage
#516Are the actual services down, or is it just the console and/or login page? For example, the sign-up page appears to be working: https://portal.aws.amazon.com/billing/signup#/start Are websites that run on AWS us-east up? Are the AWS CLIs working?
However, my Alexa (Echo) won't control my thermostat right now.
And my Ring app won't bring up my cameras.
Those services are run on AWS.
Re: AWS us-east-1 outage
#517Earlier quoted context omitted.
I worked at Walmart Technology. I bravely wrote post mortem documents owning the fault of my team (100+ people), owning both technically and also culturally as their leader. I put together a plan to fix it and executed it. Thought that was the right thing to do. This happend two times in my 10 year career there. Both times I was called out as a failure in my performance eval. Second time, I resigned and told them to…
That's shockingly stupid. I also worked for a major Walmart IT services vendor in another life, and we always had to be careful about how we handled them, because they didn't always show a lot of respect for vendors. On another note, thanks for building some awesome stuff -- walmart.com is awesome. I have both Prime and whatever-they're-currently-calling Walmart's version and I love that Walmart doesn't appear to mix…
Re: AWS us-east-1 outage
#518Friends tell friends to pick us-east-2. Virginia is for lovers, Ohio is for availability.
Re: AWS us-east-1 outage
#519Earlier quoted context omitted.
They can do this without an alliance. They very intentionally choose not to do it. Every major company has moved away from having accurate status pages.
Steam has a great status page, companies like that and cloudfare will eat Alphabet's lunch in the next 17-18 years.
Re: AWS us-east-1 outage
#520The blatant status page lies are getting absolutely ridiculous. How many hours does a service need to be totally down until it gets properly labelled as a "disruption"?