Live data from Hacker News

AWS multiple services outage in us-east-1

health.aws.amazon.com

701–710 of 1001 posts

Re: AWS multiple services outage in us-east-1

#701

Had a meeting where developers were discussing the infrastructure for an application. A crucial part of the whole flow was completely dependant on an AWS service. I asked if it was a single point of failure. The whole room laughed, I rest my case.

If you were dependent upon a single distribution (region) of that Service, yes it would be a massive single point of failure in this case. If you weren't dependent upon a particular region, you'd be fine.

Re: AWS multiple services outage in us-east-1

#702

Seems like major issues are still ongoing. If anything it seems worse than it did ~4 hours ago. For reference I'm a data engineer and it's Redshift and Airflow (AWS managed) that is FUBAR for me.

Dangerous curiosity ask, is whether the number of folks off for Diwali is a factor or not? I.e. lots of folks that weren't expected to work today and/or trying to round them up to work the problem.

Seeing as how this is us-east-1, probably not a lot.

Re: AWS multiple services outage in us-east-1

#703

Seems like major issues are still ongoing. If anything it seems worse than it did ~4 hours ago. For reference I'm a data engineer and it's Redshift and Airflow (AWS managed) that is FUBAR for me.

Definitely seems to be getting worse, outside of AWS itself, more websites seem to be having sporadic or serious issues. Concerning considering how long the outage has been going.

Re: AWS multiple services outage in us-east-1

#704

Seems like major issues are still ongoing. If anything it seems worse than it did ~4 hours ago. For reference I'm a data engineer and it's Redshift and Airflow (AWS managed) that is FUBAR for me.

Yep, confirmed worse - DynamoDB now returning "ServiceUnavailableException"

ServiceUnavailableException hello java :)

Re: AWS multiple services outage in us-east-1

#705
post #473
post #46

Choosing us-east-1 as your primary region is good, because when you're down, everybody's down, too. You don't get this luxury with other US regions!

It took me so long to realise this is what's important in enterprise. Uptime isn't important, being able to blame someone else is what's important. If you're down for 5 minutes a year because one of your employees broke something, that's your fault, and the blame passes down through the CTO. If you're down for 5 hours a year but this affected other companies too, it's not your fault From AWS to Crowdstrike - system r…

> It took me so long to realise this is what's important in enterprise. Uptime isn't important, being able to blame someone else is what's important.

Yes.

What is important is having a Contractual SLA that is defensible. Acts of God are defensible. And now major cloud infrastructure outtages are too.

Re: AWS multiple services outage in us-east-1

#708
post #687

At 3:03 AM PT AWS posted that things are recovering and sounded like issue was resolved. Then things got worse. At 9:13 AM PT it sounds like they’re back to troubleshooting. Honestly sounds like AWS doesn’t even really know what’s going on. Not good.

This is exacerbated by the fact that this is Diwali week which means the most of Indian engineers will be out on leave. Tough luck.

Re: AWS multiple services outage in us-east-1

#709

Had a meeting where developers were discussing the infrastructure for an application. A crucial part of the whole flow was completely dependant on an AWS service. I asked if it was a single point of failure. The whole room laughed, I rest my case.

Similar experience here. People laughed and some said something like "well, if something like AWS falls then we have bigger problems". They laugh because honestly is too far-fetched to think the whole AWS infra going down. Too big to fail as they say in the US. Nothing short of a nuclear war would fuck up the entire AWS network so they're kinda right.

Until this happen. A single region in a cascade failure and your saas is single region.

Re: AWS multiple services outage in us-east-1

#710

Seems like major issues are still ongoing. If anything it seems worse than it did ~4 hours ago. For reference I'm a data engineer and it's Redshift and Airflow (AWS managed) that is FUBAR for me.

Yep, confirmed worse - DynamoDB now returning "ServiceUnavailableException"

Here as well…
Post reply on HN