Live data from Hacker News

AWS us-east-1 outage

status.aws.amazon.com

741–750 of 1001 posts

Re: AWS us-east-1 outage

#741
post #372

Earlier quoted context omitted.

Yeah, but I still have a different understanding what "Increased Error Rates" means. IMHO it should mean that the rate of errors is increased but the service is still able to serve a substantial amount of traffic. If the rate of errors is bigger than, let's say, 90% that's not an increased error rate, that's an outage.

They say that to try and avoid SLA commitments.

Some big customers should get together and make an independent org to monitor cloud providers and force them to meet their SLA guarantees without being able to weasel out of the terms like this…

Re: AWS us-east-1 outage

#742

Earlier quoted context omitted.

Uh, four minutes to identify the root cause? Damn, those guys are on fire.

Identify or to publicly acknowledge? Chances are technical teams knew about this and noticed it fairly quickly, they've been working on the issue for some time. It probably wasn't until they identified the root cause and had a handful of strategies to mitigate with confidence that they chose to publicly acknowledge the issue to save face. I've broken things before and been aware of it, but didn't acknowledge them unt…

This sounds very self-blaming. Are you sure that's what's really going through your head? Personally, when I get avoidant like that, it's because of anticipation of the amount of process-related pain I'm going to have to endure as a result, and it's much easier to focus on a fix when I'm not also trying to coordinate escalation policies that I'm not familiar with.

Re: AWS us-east-1 outage

#743
post #694

I think now is a good time to reiterate the danger of companies just throwing all of their operational resilience and sustainability over the wall and trusting someone else with their entire existence. It's wild to me that so many high performing businesses simply don't have a plan for when the cloud goes down. Some of my contacts are telling me that these outages have teams of thousands of people completely prevente…

>It seems bad now but I wonder how much worse it might be when no one actually has access to money because all financial traffic is going through AWS and it goes down. Most financial institutions are implementing their own clouds, I can't think of any major one that is reliant on public cloud to the extent transactions would stop. >Why hasn't the industry come up with an alternative? You mean like building datacenter…

Indeed I can think of several outages in the past decade in the UK of banks' own infrastructure which have led to transactions stopping for days at a time, with the predictable outcomes.

Re: AWS us-east-1 outage

#744
post #245

Looks like they've acknowledged it on the status page now. https://status.aws.amazon.com/ > 8:22 AM PST We are investigating increased error rates for the AWS Management Console. > 8:26 AM PST We are experiencing API and console issues in the US-EAST-1 Region. We have identified root cause and we are actively working towards recovery. This issue is affecting the global console landing page, which is also hosted in US…

> This issue is affecting the global console landing page, which is also hosted in US-EAST-1 Even this little tidbit is a bit of a wtf for me. Why do they consider it ok to have anything hosted in a single region? At a different (unnamed) FAANG, we considered it unacceptable to have anything depend on a single region. Even the dinky little volunteer-run thing which ran https://internal.site.example/~someEngineer was…

Maybe has something to do with CloudFront mandating certs to be in us-east-1?

Re: AWS us-east-1 outage

#746

Surprised Netflix is down. I thought they were hot/hot multi-region: https://netflixtechblog.com/active-active-for-multi-regional...

Just watched two episodes of Better Call Saul on Netflix Germany without issues (while not being able to run my Terraform plan against my eu-central-1 infrastructure…).

Re: AWS us-east-1 outage

#747
post #212

A former colleague told me years ago that us-east-1 is basically the guinea pig where changes get tested before being rolled out to the other regions, and as a result is less stable than the others. Does anyone know if there's any truth to this?

At my org it was deployed in the middle, around the fourth wave iirc

Re: AWS us-east-1 outage

#748

I think now is a good time to reiterate the danger of companies just throwing all of their operational resilience and sustainability over the wall and trusting someone else with their entire existence. It's wild to me that so many high performing businesses simply don't have a plan for when the cloud goes down. Some of my contacts are telling me that these outages have teams of thousands of people completely prevente…

In my opinion there is a lack of talent in these industries for building out there own resilient systems. IT people and engineers get lazy.

Re: AWS us-east-1 outage

#749

I worked at a company that hired an ex-Amazon engineer to work on some cloud projects. Whenever his projects went down, he fought tooth and nail against any suggestion to update the status page. When forced to update the status page, he'd follow up with an extremely long "post-mortem" document that was really just a long winded explanation about why the outage was someone else's fault. He later explained that in his…

It's popular to upvote this during outages, because it fits a narrative. The truth (as always) is more complex: * No, this isn't the broad culture. It's not even a blip. These are EXCEPTIONAL circumstances by extremely bad teams that - if and when found out - would be intervened dramatically. * The broad culture is blameless post-mortems. Not whose fault is it. But what was the problem and how to fix it. And one of t…

Come on, we all know managers don’t want to claim an outage till the last minute.

Re: AWS us-east-1 outage

#750
post #245

Looks like they've acknowledged it on the status page now. https://status.aws.amazon.com/ > 8:22 AM PST We are investigating increased error rates for the AWS Management Console. > 8:26 AM PST We are experiencing API and console issues in the US-EAST-1 Region. We have identified root cause and we are actively working towards recovery. This issue is affecting the global console landing page, which is also hosted in US…

Uh, four minutes to identify the root cause? Damn, those guys are on fire.

It was down as of 7:45am (we posted in our engineering channel), so that's a good 40 minutes of public errors before the root cause was figured out.
Post reply on HN