Had a meeting where developers were discussing the infrastructure for an application. A crucial part of the whole flow was completely dependant on an AWS service. I asked if it was a single point of failure. The whole room laughed, I rest my case.
AWS multiple services outage in us-east-1
701–710 of 1001 posts
Re: AWS multiple services outage in us-east-1
#702Seems like major issues are still ongoing. If anything it seems worse than it did ~4 hours ago. For reference I'm a data engineer and it's Redshift and Airflow (AWS managed) that is FUBAR for me.
Dangerous curiosity ask, is whether the number of folks off for Diwali is a factor or not? I.e. lots of folks that weren't expected to work today and/or trying to round them up to work the problem.
Re: AWS multiple services outage in us-east-1
#703Seems like major issues are still ongoing. If anything it seems worse than it did ~4 hours ago. For reference I'm a data engineer and it's Redshift and Airflow (AWS managed) that is FUBAR for me.
Re: AWS multiple services outage in us-east-1
#704Seems like major issues are still ongoing. If anything it seems worse than it did ~4 hours ago. For reference I'm a data engineer and it's Redshift and Airflow (AWS managed) that is FUBAR for me.
Yep, confirmed worse - DynamoDB now returning "ServiceUnavailableException"
Re: AWS multiple services outage in us-east-1
#705Choosing us-east-1 as your primary region is good, because when you're down, everybody's down, too. You don't get this luxury with other US regions!
It took me so long to realise this is what's important in enterprise. Uptime isn't important, being able to blame someone else is what's important. If you're down for 5 minutes a year because one of your employees broke something, that's your fault, and the blame passes down through the CTO. If you're down for 5 hours a year but this affected other companies too, it's not your fault From AWS to Crowdstrike - system r…
Yes.
What is important is having a Contractual SLA that is defensible. Acts of God are defensible. And now major cloud infrastructure outtages are too.
Re: AWS multiple services outage in us-east-1
#706Slack was down, so I thought I will send message to my coworkers on Signal. Signal was also down.
Re: AWS multiple services outage in us-east-1
#707Whose idea was it to make the whole world dependent on us-east-1?
Re: AWS multiple services outage in us-east-1
#708At 3:03 AM PT AWS posted that things are recovering and sounded like issue was resolved. Then things got worse. At 9:13 AM PT it sounds like they’re back to troubleshooting. Honestly sounds like AWS doesn’t even really know what’s going on. Not good.
Re: AWS multiple services outage in us-east-1
#709Had a meeting where developers were discussing the infrastructure for an application. A crucial part of the whole flow was completely dependant on an AWS service. I asked if it was a single point of failure. The whole room laughed, I rest my case.
Until this happen. A single region in a cascade failure and your saas is single region.
Re: AWS multiple services outage in us-east-1
#710Seems like major issues are still ongoing. If anything it seems worse than it did ~4 hours ago. For reference I'm a data engineer and it's Redshift and Airflow (AWS managed) that is FUBAR for me.
Yep, confirmed worse - DynamoDB now returning "ServiceUnavailableException"