Live data from Hacker News

AWS multiple services outage in us-east-1

health.aws.amazon.com

961–970 of 1001 posts

Re: AWS multiple services outage in us-east-1

#962
post #941

This is having a direct impact on my wellbeing. I was at Whole Foods in Hudson Yards NYC and I couldn’t get the prime discount on my chocolate bar because the system isn’t working. Decided not to get the chocolate bar. Now my chocolate levels are way too low.

Life indeed is a struggle

First World treatlerite problems. /s What's going to suck years after too many SREs/SWEs will have long been fired, like the Morlocks & Eloi and Idiocracy, there won't be anyone left who can figure out that plants need water. There will be a few trillionaires surrounded by aristocratic, unimaginable opulence while most of humanity toils in favelas surrounded by unfixable technology that seems like magic. One cargo cult will worship 5.25" floppy disks and their arch enemies will worship CD-Rs.

https://xkcd.com/2347/

Re: AWS multiple services outage in us-east-1

#964

Looks like it affected Vercel, too. https://www.vercel-status.com/ My website is down :( (EDIT: website is back up, hooray)

I had a chuckle on my way home yesterday. Standing on the train platform and seeing "Next departure in: (Vercel Connection Error)" on the screen. :P

Re: AWS multiple services outage in us-east-1

#966
post #210

Their status page ( https://health.aws.amazon.com/health/status ) says the only disrupted service is DynamoDB, but it's impacting 37 other services. It is amazing to see how big a blast radius a single service can have.

AWS engineers are trained to use their internal services for each new system. They seem to like using DynamoDB. Dependencies like this should be made transparent.

Ex employee here who built an aws service. Dynamo is basically mandated. You need like VP approval to use a relational database because of some scaling stuff they ran into historically. That sucks because we really needed a relational database and had to bend over backwards to use dynamo and all the nonsense associated with not having sql. It was super low traffic too

Re: AWS multiple services outage in us-east-1

#967

Interesting day. I've been on an incident bridge since 3AM. Our systems have mostly recovered now with a few back office stragglers fighting for compute. The biggest miss on our side is that, although we designed a multi-region capable application, we could not run the failover process because our security org migrated us to Identity Center and only put it in us-east-1, hard locking the entire company out of the AWS…

Totally ridiculous that AWS wouldn't by default make it multi-region and warn you heavily that your multi-region service is tied to a single region for identity. The usability of AWS is so poor.

They don’t charge anything for Identity Center and so it’s not considered an important priority for the revenue counters.

Re: AWS multiple services outage in us-east-1

#968
post #395

As Amazon moves from day-1 company as it claimed once, to be the sales company like Oracle focusing on raking money, expect more outages to come, and longer to be resolved. Amazon is burning and driving away the technical talent and knowledge knowing the vendor lock-in will keep bringing the sweet money. You will see more sales people hoovering around your c-suites and executives, while you will face even worse techn…

https://www.theregister.com/2025/10/20/aws_outage_amazon_bra... quoting from this

"And so, a quiet suspicion starts to circulate: where have the senior AWS engineers who've been to this dance before gone? And the answer increasingly is that they've left the building — taking decades of hard-won institutional knowledge about how AWS's systems work at scale right along with them."

...

"AWS has given increasing levels of detail, as is their tradition, when outages strike, and as new information comes to light. Reading through it, one really gets the sense that it took them 75 minutes to go from "things are breaking" to "we've narrowed it down to a single service endpoint, but are still researching," which is something of a bitter pill to swallow. To be clear: I've seen zero signs that this stems from a lack of transparency, and every indication that they legitimately did not know what was breaking for a patently absurd length of time."

....

"This is a tipping point moment. Increasingly, it seems that the talent who understood the deep failure modes is gone. The new, leaner, presumably less expensive teams lack the institutional knowledge needed to, if not prevent these outages in the first place, significantly reduce the time to detection and recovery. "

...

"I want to be very clear on one last point. This isn't about the technology being old. It's about the people maintaining it being new. If I had to guess what happens next, the market will forgive AWS this time, but the pattern will continue."

Re: AWS multiple services outage in us-east-1

#969

Interesting day. I've been on an incident bridge since 3AM. Our systems have mostly recovered now with a few back office stragglers fighting for compute. The biggest miss on our side is that, although we designed a multi-region capable application, we could not run the failover process because our security org migrated us to Identity Center and only put it in us-east-1, hard locking the entire company out of the AWS…

"If you're able to do your job, InfoSec isn't doing theirs"

Re: AWS multiple services outage in us-east-1

#970

Interesting day. I've been on an incident bridge since 3AM. Our systems have mostly recovered now with a few back office stragglers fighting for compute. The biggest miss on our side is that, although we designed a multi-region capable application, we could not run the failover process because our security org migrated us to Identity Center and only put it in us-east-1, hard locking the entire company out of the AWS…

This reminds me of the time that Google’s Paris data center flooded and caught on fire a few years ago. We weren’t actually hosting compute there, but we were hosting compute in AWS EU datacenter nearby and it just so happened that the dns resolver for our Google services elsewhere happened to be hosted in Paris (or more accurately it routed to Paris first because it was the closest). The temp fix was pretty fun, tha…

This is the en of the thread of the first comment. Now i can find below the second comment
Post reply on HN