Live data from Hacker News

AWS multiple services outage in us-east-1

health.aws.amazon.com

991–1000 of 1001 posts

Re: AWS multiple services outage in us-east-1

#991
post #46

Choosing us-east-1 as your primary region is good, because when you're down, everybody's down, too. You don't get this luxury with other US regions!

Is us-east-1 equally unstable to the other regions? My impression was that Amazon deployed changes to us-east-1 first so it's the most unstable region.

I've heard this so many times and not seen it contradicted so I started saying it myself. Even my last Ops team wanted to run some things in us-east-1 to get prior warning before they broke us-west-1.

But there are some people on Reddit who think we are all wrong but won't say anything more. So... whatever.

Nothing in the outage history really stands out as "this is the first time we tried this and oops" except for us-east-1.

It's always possible for things to succeed at a smaller scale and fail at full scale, but again none of them really stand out as that to me. Or at least, not any in the last ten years. I'm allowing that anything older than that is on the far side of substantial process changes and isn't representative anymore.

Re: AWS multiple services outage in us-east-1

#992

Interesting day. I've been on an incident bridge since 3AM. Our systems have mostly recovered now with a few back office stragglers fighting for compute. The biggest miss on our side is that, although we designed a multi-region capable application, we could not run the failover process because our security org migrated us to Identity Center and only put it in us-east-1, hard locking the entire company out of the AWS…

Wow, you really *have* to exercise the region failover to know if it works, eh? And that confidence gets weaker the longer it’s been since the last failover I imagine too. Thanks for sharing what you learned.

The last place I worked actively switched traffic over to the backup nodes regularly (at least monthly) to ensure we could do it when necessary.

We learned that lesson by having to do emergency failovers and having some problems. :)

Re: AWS multiple services outage in us-east-1

#993
post #777

Even internal Amazon tooling is impacted greatly - including the internal ticketing platform which is making collaboration impossible during the outage. Amazon is incapable of building multi-region services internally. The Amazon retail site seems available, but I’m curious if it’s even using native AWS or is still on the old internal compute platform. Makes me wonder how much juice this company has left.

Amazon's revenue in 2024 was about the size of Belgium's GDP. Higher than Sweden or Ireland. It makes a profit similar to Norway, without drilling for offshore oil or maintaining a navy. I think they've got plenty of juice left.

The universe's metaphysical poeticism holds that it's slightly more likely than it otherwise would be that the company that replaced Sears would one day go the way of Sears.

Re: AWS multiple services outage in us-east-1

#995
We never went down in us-east-1 during this incident. We have tons of high-traffic sites/services. Not multi-region, not multi-cloud.

You're gonna hear mostly complaints in this thread, but simple, resilient, single-region architecture is still reliable as hell in AWS, even in the worst region.

Re: AWS multiple services outage in us-east-1

#996
post #805

Seems like major issues are still ongoing. If anything it seems worse than it did ~4 hours ago. For reference I'm a data engineer and it's Redshift and Airflow (AWS managed) that is FUBAR for me.

Andy Jassy is the Tim Cook of Amazon Rest and vest CEOs

Don’t insult Tim Cook like that.

He got a lot of impossible shit done as COO.

They do need a more product minded person though. If Jobs was still around we’d have smart jewelry by now. And the Apple Watch would be thin af.

Re: AWS multiple services outage in us-east-1

#997

Seems like major issues are still ongoing. If anything it seems worse than it did ~4 hours ago. For reference I'm a data engineer and it's Redshift and Airflow (AWS managed) that is FUBAR for me.

The problem is now that, what’s anyone going to do? Leave? I remember a meme years ago about Nestle. It was something like: GO ON, BOYCOT US - I BET YOU CAN’T - WE MAKE EVERYTHING. Same meme would work for Aws today.

It’s amazing how much you can avoid them by eating food that still looks like what it started as though. They own a lot of processed food.

Re: AWS multiple services outage in us-east-1

#998

Seems like major issues are still ongoing. If anything it seems worse than it did ~4 hours ago. For reference I'm a data engineer and it's Redshift and Airflow (AWS managed) that is FUBAR for me.

Agreed, every time the impacted services list internally gets shorter, the next update it starts growing again. A lot of these are second order dependencies like Astronomer, Atlassian, Confluent, Snowflake, Datadog, etc... the joys of using hosted solutions to everything.

Before my old company spun off, we didn’t know the old ops team had put on-prem production and our Atlassian instances in the same NAS.

When the NAS shit the bed, we lost half of production and all our run books. And we didn’t have autoscaling yet. Wouldn’t for another 2 years.

Our group is a bunch of people that has no problem getting angry and raising voices. The whole team was so volcanically angry that it got real quiet for several days. Like everyone knew if anyone unclenched that there would be assault charges.

Re: AWS multiple services outage in us-east-1

#999
post #809

Seems like major issues are still ongoing. If anything it seems worse than it did ~4 hours ago. For reference I'm a data engineer and it's Redshift and Airflow (AWS managed) that is FUBAR for me.

This looks like one their worst outage in 15 years and us-east-1 still shows as degraded but I had no outages, as dont use us-east-1. Are you seeing issues on other regions? https://health.aws.amazon.com/health/status?path=open-issues The closest to their identification of a root cause seems to be this one: "Oct 20 8:43 AM PDT We have narrowed down the source of the network connectivity issues that impacted AWS Servi…

I wonder how many people discovered their autoscaling settings went batshit when services went offline, either scaling way down or way up, or went metastable and started fishtailing.

Re: AWS multiple services outage in us-east-1

#1000

Seems like major issues are still ongoing. If anything it seems worse than it did ~4 hours ago. For reference I'm a data engineer and it's Redshift and Airflow (AWS managed) that is FUBAR for me.

Dangerous curiosity ask, is whether the number of folks off for Diwali is a factor or not? I.e. lots of folks that weren't expected to work today and/or trying to round them up to work the problem.

Sometimes I miss my phone buzzing when doing yard work. Diwali has to be worse for that.
Post reply on HN