Live data from Hacker News

AWS multiple services outage in us-east-1

health.aws.amazon.com

651–660 of 1001 posts

Re: AWS multiple services outage in us-east-1

#651
Paying for resilience is expensive. not as expensive as AWS, but it's not free.

Modern companies live life on the edge. Just in time, no resilience, no flexibility. We see the disaster this causes whenever something unexpected happens - the Evergiven blocking Suez for example, let alone something like Covid

However increasingly what should be minor loss of resilience, like an AWS outage or a Crowdstrike incident, turns into major failures.

This fragility is something government needs to legislate to prevent. When one supermarket is out that's fine - people can go elsewhere, the damage is contained. When all fail, that's a major problem.

On top of that, the attitude that the entire sector has is also bad. People thing IT should tail once or twice a year and it's not a problem. If that attitude affect truly important systems it will lead to major civil projects. Any civilitsation is 3 good meals away from anarchy.

There's no profit motive to avoid this, companies don't care about being offline for the day, as long as all their mates are also offline.

Re: AWS multiple services outage in us-east-1

#652

I forget where I read it originally, but I strongly feel that AWS should offer a `us-chaos-1` region, where every 3-4 days, one or two services blow up. Host your staging stack there and you build real resiliency over time. (The counter joke is, of course, "but that's `us-east-1` already! But I mean deliberately and frequently.)

AWS already offers Fault Injection Service, which you can use in any region to conduct chaos engineering: https://aws.amazon.com/fis/

Re: AWS multiple services outage in us-east-1

#653
post #610

This is just a silly anecdote, but every time a cloud provider blips, I'm reminded. The worst architecture I've ever encountered was a system that was distributed across AWS, Azure, and GCP. Whenever any one of them had a problem, the system went down. It also cost 3x more than it should.

I've seen the exact same thing at multiple companies. The teams were always so proud of themselves for being "multi-cloud" and managers rewarded them for their nonsense. They also got constant kudos for their heroic firefighting whenever the system went down, which it did constantly. Watching actually good engineers get overlooked because their systems were rock-solid while those characters got all the praise for des…

Were you able to find a better company? If yes, what kind of company?

Re: AWS multiple services outage in us-east-1

#654
I in-housed an EMR for a local clinic because of latency and other network issues taking the system offline several times a month (usually at least once a week). We had zero downtime the whole first year after bringing it all in house, and I got employee of the month for several months in a row.

Re: AWS multiple services outage in us-east-1

#655

This is just a silly anecdote, but every time a cloud provider blips, I'm reminded. The worst architecture I've ever encountered was a system that was distributed across AWS, Azure, and GCP. Whenever any one of them had a problem, the system went down. It also cost 3x more than it should.

looks like very few get it right. A good system would have few minutes of blip when one cloud provider went down, which is a massive win compared to outages like this.

They all make it pretty hard, and a lot of resume-driven-devs have a hard time resisting the temptation of all the AWS alphabet soup of services.

Sure you can abstract everything away, but you can also just not use vendor-flavored services. The more bespoke stuff you use the more lock in risk.

But if you are in a "cloud forward" AWS mandated org, a holder of AWS certifications, alphabet soup expert... thats not a problem you are trying to solve. Arguably the lock in becomes a feature.

Re: AWS multiple services outage in us-east-1

#656

Interesting day. I've been on an incident bridge since 3AM. Our systems have mostly recovered now with a few back office stragglers fighting for compute. The biggest miss on our side is that, although we designed a multi-region capable application, we could not run the failover process because our security org migrated us to Identity Center and only put it in us-east-1, hard locking the entire company out of the AWS…

I remember Facebook had a similar story when they botched their BGP update and couldn't even access the vault. If you have circular auth, you don't have anything when somebody breaks DNS.

Wasn't there an issue where they required physical access to the data center to fix the network, which meant having to tap in with a keycard to get in, which didn't work because the keycard server was down, due to the network being down?

Re: AWS multiple services outage in us-east-1

#658
Wow, about 9 hours later and 21 of 24 Atlassian services are still showing up as impacted on their status page.

Even @ 9:30am ET this morning, after this supposedly was clearing up, my doctor's office's practice management software was still hosed. Quite the long tail here.

https://status.atlassian.com/

Re: AWS multiple services outage in us-east-1

#659
post #609

Er...They appear to have just gone down again.

My systems didn't actually seem to be affected until what I think was probably a SECOND spike of outages at about the time you posted.

The retrospective will be very interesting reading!

(Obviously the category of outages caused by many restored systems "thundering" at once to get back up is known, so that'd be my guess, but the details are always good reading either way).

Re: AWS multiple services outage in us-east-1

#660

A lot of status pages hosted by Atlasian StatusPage are down! The irony…

I can't believe this. When status page first created their product, they used to market how they were in multiple providers so that they'd never be affected by downtime.

Maybe all that got canned after the acquisition?

Post reply on HN