Live data from Hacker News

AWS multiple services outage in us-east-1

health.aws.amazon.com

731–740 of 1001 posts

Re: AWS multiple services outage in us-east-1

#731

Interesting day. I've been on an incident bridge since 3AM. Our systems have mostly recovered now with a few back office stragglers fighting for compute. The biggest miss on our side is that, although we designed a multi-region capable application, we could not run the failover process because our security org migrated us to Identity Center and only put it in us-east-1, hard locking the entire company out of the AWS…

> Identity Center and only put it in us-east-1 Is it possible to have it in multiple regions? Last I checked, it only accepted one region. You needed to remove it first if you wanted to move it.

Correct. That does make it a centralized failure mode and everyone is in the same boat on that.

I’m unaware of any common and popular distributed IDAM that is reliable

Re: AWS multiple services outage in us-east-1

#732

Seems like major issues are still ongoing. If anything it seems worse than it did ~4 hours ago. For reference I'm a data engineer and it's Redshift and Airflow (AWS managed) that is FUBAR for me.

Downdetector had 5,755 reports of AWS problems at 12:52 AM Pacific (3:53 AM Eastern). That number had dropped to 1,190 by 4:22 AM Pacific (7:22 AM Eastern). However, that number is back up with a vengeance. 9,230 reports as of 9:32 AM Pacific (12:32 Eastern). Part of that could be explained by more people making reports as the U.S. west coast awoke. But I also have a feeling that they aren't yet on top of the problem…

Where do they source those reports from? Always wondered if it was just analysis of how many people are looking at the page, or if humans somewhere are actually submitting reports.

Re: AWS multiple services outage in us-east-1

#733
post #610

This is just a silly anecdote, but every time a cloud provider blips, I'm reminded. The worst architecture I've ever encountered was a system that was distributed across AWS, Azure, and GCP. Whenever any one of them had a problem, the system went down. It also cost 3x more than it should.

I've seen the exact same thing at multiple companies. The teams were always so proud of themselves for being "multi-cloud" and managers rewarded them for their nonsense. They also got constant kudos for their heroic firefighting whenever the system went down, which it did constantly. Watching actually good engineers get overlooked because their systems were rock-solid while those characters got all the praise for des…

multi-cloud... any leader that approves such a boondoggle should be labelled incompetent. These morons sell it as a cost-cutting "migration". Never once have I seen such a project complete and it more than doubles complexity and costs.

Re: AWS multiple services outage in us-east-1

#737
post #687

At 3:03 AM PT AWS posted that things are recovering and sounded like issue was resolved. Then things got worse. At 9:13 AM PT it sounds like they’re back to troubleshooting. Honestly sounds like AWS doesn’t even really know what’s going on. Not good.

This is exacerbated by the fact that this is Diwali week which means the most of Indian engineers will be out on leave. Tough luck.

[flagged]

Re: AWS multiple services outage in us-east-1

#739

Seems like major issues are still ongoing. If anything it seems worse than it did ~4 hours ago. For reference I'm a data engineer and it's Redshift and Airflow (AWS managed) that is FUBAR for me.

I'm wondering why your and other companies haven't just evicted themselves from us-east-1. It's the worst region for outages and it's not even close. Our company decided years ago to use any region other than us-east-1. Of course, that doesn't help with services that are 'global', which usually means us-east-1.

Some AWS services are only available in us-east-1. Also a lot of people have not built their infra to be portable and the occasional outage isn't worth the cost and effort of moving out.
Post reply on HN