Live data from Hacker News

AWS us-east-1 outage

status.aws.amazon.com

281–290 of 1001 posts

Re: AWS us-east-1 outage

#281
post #28

Earlier quoted context omitted.

Status pages are hard

When they have too much pride in an all-green dash, sure. Allowing any engineer to declare a problem when first detected? Not so hard, but it doesn't make you look good if you have an ultra-twitchy finger. They have the balance badly wrong at the moment though.

A trigger-happy status page gives realtime feedback for anyone doing a DoS attack. Even if you published that information publicly you would probably want it on a significant delay.

Re: AWS us-east-1 outage

#284
post #217

Earlier quoted context omitted.

That sounds like the exact opposite of human-factors engineering. No one likes taking blame. But when things go sideways, people are extra spicy and defensive, which makes them clam up and often withhold useful information, which can extend the outage. No-blame analysis is a much better pattern. Everyone wins. It's about building the system that builds the system. Stuff broke; fix the stuff that broke, then fix the t…

> No-blame analysis is a much better pattern. Everyone wins. It's about building the system that builds the system. Stuff broke; fix the stuff that broke, then fix the things that let stuff break. Yea, except it doesn't work in practice. I work with a lot of people who come from places with "blameless" post-mortem 'culture' and they've evangelized such a thing extensively. You know what all those people have proven t…

Ok, and? I don't doubt it fails in places. That doesn't mean that it doesn't work in practice. Our company does it just fine. We have a high trust, high transparency system and it's wonderful.

It's like saying unit tests don't work in practice because bugs got through.

Re: AWS us-east-1 outage

#285
post #245

Looks like they've acknowledged it on the status page now. https://status.aws.amazon.com/ > 8:22 AM PST We are investigating increased error rates for the AWS Management Console. > 8:26 AM PST We are experiencing API and console issues in the US-EAST-1 Region. We have identified root cause and we are actively working towards recovery. This issue is affecting the global console landing page, which is also hosted in US…

They are still lying about it, the issues are not only affecting the console but also AWS operations such as S3 puts. S3 still shows green.

It's certainly affecting a wider range of stuff from what I've seen. I'm personally having issues with API Gateway, CloudFormation, S3, and SQS

Re: AWS us-east-1 outage

#286

It's funny that the first place I go to learn about the outage is Hacker News and not https://status.aws.amazon.com/ (it's still reports everything to be "operating normally"...)

I made sure our incident response plan includes checking Hacker News and Twitter for actual updates and information. As of right now, this thread and one update from a twitter user, https://twitter.com/SiteRelEnby/status/1468253604876333059 are all we have. I went into disaster recovery mode when I saw our traffic dropped to 0 suddenly at 10:30am ET. That was just the SQS/something else preventing our ELB logs from b…

So as of the time you posted this comment, were other services actually down? The way the 500 shows up, and the AWS status page, makes it sound like "only" the main landing page/mgt console is unavailable, not AWS services.

Re: AWS us-east-1 outage

#289
post #70

Earlier quoted context omitted.

Especially if enough Amazon internal tools rely on it - would be funny if there were a repeat of the FB debacle where Amazon employees somehow couldn't communicate/get back into their offices because of the problem they were trying to fix

Last I knew, Amazon used all Microsoft stuff for business communication.

Slack, as of last year.

https://slack.com/blog/news/slack-aws-drive-development-agil...

Re: AWS us-east-1 outage

#290
post #212

A former colleague told me years ago that us-east-1 is basically the guinea pig where changes get tested before being rolled out to the other regions, and as a result is less stable than the others. Does anyone know if there's any truth to this?

false, it's often 4th iirc, SFO (us-west-1) is actually usually first.
Post reply on HN