https://status.aws.amazon.com hasn't been updated to reflect outage yest
AWS us-east-1 outage
541–550 of 1001 posts
Re: AWS us-east-1 outage
#542My job (although 50% of time) at Azure is unit testing/monitoring services under different scenarios and flows to detect small failures that will be overlooked in public status page. Our tests run multiple times daily and we have people constantly monitoring logs. It concerns me when I see all AWS services are 100% green when I know there is an outage.
The reason you care about your status page being 100% accurate is that your stock price is not directly linked to your status page.
Re: AWS us-east-1 outage
#543I'm not complaining, I enjoyed the nostaliga - sometimes the web still feels like the late 90s
Re: AWS us-east-1 outage
#544Earlier quoted context omitted.
Well, the narrative is sort of what Amazon is asking for, heh? The whole us-east-1 management console is gone, what is Amazon posting for the management console on their website? "Service degradation" It's not a degradation if it's outright down. Use the red status a little bit more often, this is a "disruption", not a "degradation".
I've always wondered why services are not counted down more often. Is there some sliver of customers who have access to the management console for example? An increase in error rates - no biggie, any large system is going to have errors. But when 80%+ of customers loads in the region are impacted (cross availability zones for whatever good those do) - that counts as down doesn't it? Error rates in one AZ - degraded.…
Re: AWS us-east-1 outage
#545Looks like they've acknowledged it on the status page now. https://status.aws.amazon.com/ > 8:22 AM PST We are investigating increased error rates for the AWS Management Console. > 8:26 AM PST We are experiencing API and console issues in the US-EAST-1 Region. We have identified root cause and we are actively working towards recovery. This issue is affecting the global console landing page, which is also hosted in US…
Uh, four minutes to identify the root cause? Damn, those guys are on fire.
> Dev1: Pushing code for branch "master" to "AWS API". > Your deploy finished in 4 minutes > Dev2: I can't react the API in east-1 > Dev1: Works from my computer
Re: AWS us-east-1 outage
#546Re: AWS us-east-1 outage
#547I was currently in the process of buying some tickets on Ticketmaster and the entire presale event had to be postponed for at least 4 hours due to this AWS outage. I'm not complaining, I enjoyed the nostaliga - sometimes the web still feels like the late 90s
Re: AWS us-east-1 outage
#548Earlier quoted context omitted.
So as of the time you posted this comment, were other services actually down? The way the 500 shows up, and the AWS status page, makes it sound like "only" the main landing page/mgt console is unavailable, not AWS services.
Yes, they are still publishing lies on their status page. In this thread people are reporting issues with many services. I'm seeing periodic S3 PUT failures for the last 1.5 hours.
Re: AWS us-east-1 outage
#549Earlier quoted context omitted.
If you're not multi-cloud in 2021 and are expecting 5-9's, I feel bad for you.
If you're not multi-region, I feel bad for you. If your company is shoehorning you into using multiple clouds and learning a dozen products, IAM and CICD dialects simultaneously because "being cloud dependent is bad", I feel bad for you. Doing one cloud correctly from a current DevSecOps perspective is a multi-year ask. I estimate it takes about 25 people working full time on managing and securing infrastructure per…
Example: Payment/Administrative issues, rogue employee with access, deprecated service, inter-region routing issues, root certificate compromises... the list goes on and it is certainly not limited to single AZ.
A very good example, is that regardless of which of the 85 AZs you are in at aws, you are affected by this issue right now.
Multi-cloud with the right tooling is trivial. Investing in learning cloud-proprietary stacks is a waste of your investment. You're a clown if you think you need 25 people internally per cloud is required to "do it right".
Re: AWS us-east-1 outage
#550Earlier quoted context omitted.
That sounds like the exact opposite of human-factors engineering. No one likes taking blame. But when things go sideways, people are extra spicy and defensive, which makes them clam up and often withhold useful information, which can extend the outage. No-blame analysis is a much better pattern. Everyone wins. It's about building the system that builds the system. Stuff broke; fix the stuff that broke, then fix the t…
I don't think engineers can believe in no-blame analysis if they know it'll harm career growth. I can't unilaterally promote John Doe, I have to convince other leaders that John would do well the next level up. And in those discussions, they could bring up "but John has caused 3 incidents this year", and honestly, maybe they'd be right.