Live data from Hacker News

AWS multiple services outage in us-east-1

health.aws.amazon.com

591–600 of 1001 posts

Re: AWS multiple services outage in us-east-1

#592
post #487

US-East-1 is more than just a normal region. It also provides the backbone for other services, including those in other regions. Thus simply being in another region doesn’t protect you from the consistent us-east-1 shenanigans. AWS doesn’t talk about that much publicly, but if you press them they will admit in private that there are some pretty nasty single points of failure in the design of AWS that can materialize…

Amazon are planning to launch the EU Sovereign Cloud by the end of the year. They claim it will be completely independent. It may be possible then to have genuine resiliency on AWS. We'll see.

Re: AWS multiple services outage in us-east-1

#593

If we see more of this, it would not be crazy to assume that all this compelling of engineers to "use AI" and the flood of Looks Good To Me code is coming home.

Big if, major outages like this aren't unheard of, and so far, fairly uncommon. Definitely hit harder than their SLAs promise though. I hope they do an honest postmortem, but I doubt they would blame AI even if it was somehow involved. Not to mention you can't blame AI unless you go completely hands-off - but that's like blaming an outsourcing partner, which also never happens.

Re: AWS multiple services outage in us-east-1

#594
post #473
post #46

Choosing us-east-1 as your primary region is good, because when you're down, everybody's down, too. You don't get this luxury with other US regions!

It took me so long to realise this is what's important in enterprise. Uptime isn't important, being able to blame someone else is what's important. If you're down for 5 minutes a year because one of your employees broke something, that's your fault, and the blame passes down through the CTO. If you're down for 5 hours a year but this affected other companies too, it's not your fault From AWS to Crowdstrike - system r…

A slightly less cynical view: execs have a hard filter for “things I can do something about” and “things I can’t influence at all.” The bad ones are constantly pushing problems into the second bucket, but there are legitimately gray area cases. When an exec smells the possibility that their team could have somehow avoided a problem, that’s category 1 and the hammer comes down hard.

After that process comes the BS and PR step, where reality is spun into a cotton candy that makes the leader look good no matter what.

Re: AWS multiple services outage in us-east-1

#596

One thing has become quite clear to me over the years. Much of the thinking around uptime of information systems has become hyperbolic and self-serving. There are very few businesses that genuinely cannot handle an outage like this. The only examples I've personally experienced are payment processing and semiconductor manufacturing. A severe IT outage in either of these businesses is an actual crisis. Contrast with t…

Tech companies, and in particular ad-driven companies, keep a very close eye on their metrics and can fairly accurately measure the cost of an outage in real dollars

Re: AWS multiple services outage in us-east-1

#598

stupid question: is buying a server rack and running it at home subject to more downtimes in a year than this? has anyone done an actual SLA analysis?

That depends on a lot of factors, but for me personally, yes it is. Much worse.

Assuming we’re talking about hosting things for Internet users. My fiber internet connection has gone down multiple times, though relatively quickly restored. My power has gone out several times in the last year, with one storm having it out for nearly 24 hrs. I was sleep when it went out and I didn’t start the generator until it was out for 3-4 hours already, far longer than my UPSes could hold up. I’ve had to do maintenance and updates both physical and software.

All of those things contribute to a downtime significantly higher than I see with my stuff running on Linode, Fly.io or AWS.

I run Proxmox and K3s at home and it makes things far more reliable, but it’s also extra overhead for me to maintain.

Most or all of those things could be mitigated at home, but at what cost?

Post reply on HN