Live data from Hacker News

AWS multiple services outage in us-east-1

health.aws.amazon.com

901–910 of 1001 posts

Re: AWS multiple services outage in us-east-1

#901

Interesting day. I've been on an incident bridge since 3AM. Our systems have mostly recovered now with a few back office stragglers fighting for compute. The biggest miss on our side is that, although we designed a multi-region capable application, we could not run the failover process because our security org migrated us to Identity Center and only put it in us-east-1, hard locking the entire company out of the AWS…

Too much armor makes you immobile. Will your security org be held to task for this? This should permanently slow down all of their future initiatives because it’s clear they have been running “faster than possible” for some time.

Who watches the watchers.

Re: AWS multiple services outage in us-east-1

#902

Seems like major issues are still ongoing. If anything it seems worse than it did ~4 hours ago. For reference I'm a data engineer and it's Redshift and Airflow (AWS managed) that is FUBAR for me.

It has been quite a while, wondering how many 9s are dropped. 365 day * 24 * 0.0001 is roughly 8 hours, so it already lost the 99.99% status.

Where were you guys the other day when someone was calling me crazy for trying to make this same sort of argument?

Re: AWS multiple services outage in us-east-1

#903
I missed a parcel delivery because a computer server in Virginia, USA went down, and now the doorbell on my house in England doesn't work. What. The. Fork.

How the hell did Ring/Amazon not include a radio-frequency transmitter for the doorbell and chime? This is absurd.

To top it off, I'm trying to do my quarterly VAT return, and Xero is still completely borked, nearly 20 hours after the initial outage.

Re: AWS multiple services outage in us-east-1

#904

Seems like major issues are still ongoing. If anything it seems worse than it did ~4 hours ago. For reference I'm a data engineer and it's Redshift and Airflow (AWS managed) that is FUBAR for me.

Dangerous curiosity ask, is whether the number of folks off for Diwali is a factor or not? I.e. lots of folks that weren't expected to work today and/or trying to round them up to work the problem.

Seems like a lot of people missing that this post was made around midnight PST time and thus it would be more reasonable to ping people at lunch in IST before waking up people in EST or PST.

Re: AWS multiple services outage in us-east-1

#905

Seems like major issues are still ongoing. If anything it seems worse than it did ~4 hours ago. For reference I'm a data engineer and it's Redshift and Airflow (AWS managed) that is FUBAR for me.

You have to remember that health status dashboards at most (all?) cloud providers require VP approval to switch status. This stuff is not your startup's automated status dashboard. It's politics, contracts, money.

Which makes them a flat out lie since it ceases to be a dashboard if it’s not live. It’s just a status page.

Re: AWS multiple services outage in us-east-1

#906

The Premier League said there will be only limited VAR today w/o the automatic offside system becasue of the AWS outage. Weird timeline we live in https://www.bbc.com/news/live/c5y8k7k6v1rt?post=asset%3Ad902...

A silver lining to this cloud (outage).

Re: AWS multiple services outage in us-east-1

#907
post #777

Even internal Amazon tooling is impacted greatly - including the internal ticketing platform which is making collaboration impossible during the outage. Amazon is incapable of building multi-region services internally. The Amazon retail site seems available, but I’m curious if it’s even using native AWS or is still on the old internal compute platform. Makes me wonder how much juice this company has left.

Amazon's revenue in 2024 was about the size of Belgium's GDP. Higher than Sweden or Ireland. It makes a profit similar to Norway, without drilling for offshore oil or maintaining a navy. I think they've got plenty of juice left.

The navy comment is a bit unfair, as it's well-known that Amazon is more of an airpower (hence "the cloud" etc.)

Re: AWS multiple services outage in us-east-1

#908

The Premier League said there will be only limited VAR today w/o the automatic offside system becasue of the AWS outage. Weird timeline we live in https://www.bbc.com/news/live/c5y8k7k6v1rt?post=asset%3Ad902...

Why is VAR connected to the internet? Are they trying to gather data on offside players customers to improve recommedations?

Re: AWS multiple services outage in us-east-1

#909

Interesting day. I've been on an incident bridge since 3AM. Our systems have mostly recovered now with a few back office stragglers fighting for compute. The biggest miss on our side is that, although we designed a multi-region capable application, we could not run the failover process because our security org migrated us to Identity Center and only put it in us-east-1, hard locking the entire company out of the AWS…

I remember Facebook had a similar story when they botched their BGP update and couldn't even access the vault. If you have circular auth, you don't have anything when somebody breaks DNS.

That's similar to the total outage of all Rogers services in Canada back on July 7th 2022. It was compounded by the fact that the outage took out all Rogers cell phone service, making it impossible for Rogers employees to communicate with each other during the outage. A unified network means a unified failure mode.

Thankfully none of my 10 Gbps wavelengths were impacted. Oh did I appreciate my aversion to >= layer 2 services in my transport network!

Re: AWS multiple services outage in us-east-1

#910

Their status page ( https://health.aws.amazon.com/health/status ) says the only disrupted service is DynamoDB, but it's impacting 37 other services. It is amazing to see how big a blast radius a single service can have.

i'm surprised / bothered that the history log shows the issues starting AM 10/20 -- when they seemed to have started around midnight 10/19
Post reply on HN