Live data from Hacker News

AWS us-east-1 outage

status.aws.amazon.com

311–320 of 1001 posts

Re: AWS us-east-1 outage

#311
post #66

Are the actual services down, or is it just the console and/or login page? For example, the sign-up page appears to be working: https://portal.aws.amazon.com/billing/signup#/start Are websites that run on AWS us-east up? Are the AWS CLIs working?

I wasn't able to load my amazon.com wishlist, nor the shopping page through the app. Not an aws service specifically, but an amazon service that I couldn't use.

Re: AWS us-east-1 outage

#312
post #113

Friends tell friends to pick us-east-2. Virginia is for lovers, Ohio is for availability.

I live in Ohio and can confirm. If the Earth were destroyed by an asteroid Ohio would be left floating out there somehow holding onto an atmosphere for about ten years.

Re: AWS us-east-1 outage

#313

After over 45 minutes https://status.aws.amazon.com/ now shows "AWS Management Console - Increased Error Rates" I guess 100% is technically an increase.

"Fixed a bug that could cause [adverse behavior affecting 100% of the user base] for some users"

Re: AWS us-east-1 outage

#314
post #294

I worked at a company that hired an ex-Amazon engineer to work on some cloud projects. Whenever his projects went down, he fought tooth and nail against any suggestion to update the status page. When forced to update the status page, he'd follow up with an extremely long "post-mortem" document that was really just a long winded explanation about why the outage was someone else's fault. He later explained that in his…

This is the exact opposite of my experience at AWS. Amazon is all about blameless fact finding when it comes to root cause analysis. Your company just hired a not so great engineer or misunderstood him.

Adding my piece of anecdata to this.. the process is quite blameless. If a postmortem seems like it points blame, this is pointed out and removed.

Re: AWS us-east-1 outage

#315
post #217

Earlier quoted context omitted.

That sounds like the exact opposite of human-factors engineering. No one likes taking blame. But when things go sideways, people are extra spicy and defensive, which makes them clam up and often withhold useful information, which can extend the outage. No-blame analysis is a much better pattern. Everyone wins. It's about building the system that builds the system. Stuff broke; fix the stuff that broke, then fix the t…

I firmly believe in the dictum "if you ship it you own it". That means you own all outages. It's not just an operator flubbing a command, or a bit of code that passed review when it shouldn't. It's all your dependencies that make your service work. You own ALL of them. People spend all this time threat modelling their stuff against malefactors, and yet so often people don't spend any time thinking about the threat mo…

when working on CloudFiles, we often had monitoring for our limited dependencies that were better than their monitoring. Don't just know what your stuff is doing, but what your whole dependency ecosystem is doing and know when it all goes south. also helps to learn where and how you can mitigate some of those dependencies.

Re: AWS us-east-1 outage

#316

This got me thinking, are there any major chat services that would go down if a particular AWS/GCP/etc data centre went down? You don't want your service to go down, plus your team's comms at the same time.

We have a SMS text thread with about 12 people that we send one message on the first of every month. To make sure it is tested and ready to be used for communications if all other comms networks are down.

Re: AWS us-east-1 outage

#317

Earlier quoted context omitted.

Last I knew, Amazon used all Microsoft stuff for business communication.

Slack, as of last year. https://slack.com/blog/news/slack-aws-drive-development-agil...

And before that, Amazon Chime was the messaging and conferencing tool. Now that I'm not using it, I actually miss it a lot!

Re: AWS us-east-1 outage

#318

After over 45 minutes https://status.aws.amazon.com/ now shows "AWS Management Console - Increased Error Rates" I guess 100% is technically an increase.

I can't remember seeing problems be more strongly worded than "Increased Error Rates" or "high error rates with S3 in us-east-1" during the infamous S3 outage of 2017 - and that was after they struggled to even update their own status page because of S3 being down. :)

Re: AWS us-east-1 outage

#319

AWS Connect is down, so our customer support phone system is down with it

Highly recommend talking to your account team to recommend regional failovers and DR for Amazon Connect! With enough feedback from customers, stuff like this can get prioritized.

Re: AWS us-east-1 outage

#320

Earlier quoted context omitted.

They are still lying about it, the issues are not only affecting the console but also AWS operations such as S3 puts. S3 still shows green.

It's certainly affecting a wider range of stuff from what I've seen. I'm personally having issues with API Gateway, CloudFormation, S3, and SQS

> We are experiencing API and console issues in the US-EAST-1 Region
Post reply on HN