Live data from Hacker News

AWS us-east-1 outage

status.aws.amazon.com

971–980 of 1001 posts

Re: AWS us-east-1 outage

#971
post #759

Earlier quoted context omitted.

So you're saying companies should start moving their infrastructure to the blockchain?

Ethereum has gone 5 years without a single minute of downtime, so if it's extreme reliability you're going for I don't think it can be beaten.

Easy to be reliable when nobody uses it

Re: AWS us-east-1 outage

#972

Earlier quoted context omitted.

Postmortem: unbounded auto-scaling of lambda combined with oversight on internal rate limits caused unforseen internal ddos.

Just wait for the medium article “How I ran up a $400 million AWS bill.”

"On how I learned 'recursion'"

Re: AWS us-east-1 outage

#973
post #897

Earlier quoted context omitted.

Pretty low chance that the status page is automated, especially via health checks. I imagine it's a static asset updated by hand.

Or the service that updates the status page runs out of us-east-1.

It has customer relationship implications. I guarantee you it is updated by a support agent.

Re: AWS us-east-1 outage

#975

Earlier quoted context omitted.

So basically it's just a toggle switch?

My lighting automation uses Insteon currently. My primary req is that they are smart and connected without needing a central controller o a connection to the internet. My switches all understand lighting scenes and can manage those in a P2P manner, without a central controller. The central controller is primarily used when I want to add actual automations vs. scenes. Even the central controller aspect works 100% disc…

[deleted]

Re: AWS us-east-1 outage

#976
post #960

Earlier quoted context omitted.

Thanks. That shows that OPs claim that "Usually, if us-east-1 goes down, others soon follow." is false.

Cognito (and r53) have hard dependencies in us-east-1. It’s mentioned up and down this thread.

I don't want to minimize the impact from cognito and r53, but that's quite a different scale of failure than the OP implies. It's been a while since I used AWS, but we had multiple regions and saw no impact to our services in other regions the one time that us-east-1 had a major failure. And we used r53.

Re: AWS us-east-1 outage

#977
post #960

Earlier quoted context omitted.

Cognito (and r53) have hard dependencies in us-east-1. It’s mentioned up and down this thread.

I don't want to minimize the impact from cognito and r53, but that's quite a different scale of failure than the OP implies. It's been a while since I used AWS, but we had multiple regions and saw no impact to our services in other regions the one time that us-east-1 had a major failure. And we used r53.

Perhaps I could've been more precise with my words. What I meant to say is IAM and r53 are two of many critical services that all depend on us-east-1. It goes without saying that if those services go down in us-east-1, the whole AWS is affected. This doesn't just happen "usually." When IAM goes down, AWS experiences major issues across all regions. If you were okay, perhaps you got lucky? Our team has 7 different prod regions and we see multiple regions go down every time a problem of this scale occurs.

If your product requires 100% uptime, you may need to look at backup options or design your product in such a way that can handle temporary cloud failures.

Re: AWS us-east-1 outage

#978
post #13

I love that every time this happens, 100% of the services on https://status.aws.amazon.com are green.

I don't see why they couldn't provide an error rate graph like Reddit[0] or simply make services yellow saying "increased error rate detected, investigating..." 0: https://www.redditstatus.com/#system-metrics

Wow. Kudos to the reddit engineering team. That's one of the nicest status pages I have seen.

Re: AWS us-east-1 outage

#979

Earlier quoted context omitted.

YES! Why do they do that? It's so weird. I will deploy a whole config into us-west-1 or something; but then I need to create a new cert in us-east-1 JUST to let cloudfront answer an HTTPS call. So frustrating.

Agreed - in my line of work regulators want everything in the country we operate from but of course CloudFront has to be different.

Wouldn't using a global CDN for everything be off the table to begin with, in that case?

Re: AWS us-east-1 outage

#980

Earlier quoted context omitted.

It’s standard. Career ladder [1] sets expectation for each level. Performance is measured against those expectations. Outages don’t negatively impact a single engineer. The key difference is the perspective. If reliability is bad that’s an organizational problem and blaming or punishing one engineer won’t fix that. [1] An example ladder from Patreon: https://levels.patreon.com/

> The key difference The key difference between what and what?

[deleted]
Post reply on HN