Live data from Hacker News

AWS us-east-1 outage

status.aws.amazon.com

251–260 of 1001 posts

Re: AWS us-east-1 outage

#251
post #66

Are the actual services down, or is it just the console and/or login page? For example, the sign-up page appears to be working: https://portal.aws.amazon.com/billing/signup#/start Are websites that run on AWS us-east up? Are the AWS CLIs working?

One of my sites went offline an hour ago because the web server stopped responding. I can't SSH into it or get any type of response. The database server in the same region and zone is continuing to run fine though.

Re: AWS us-east-1 outage

#252

I can still hit EC2 boxes and networking is okay. DynamoDB is 100% down for the count, every request is an Internal Server Error.

DynamoDB is fine for us. Not contradicting your experience, just adding another data point. There is definitely something hit-or-miss about this incident.

Re: AWS us-east-1 outage

#253

At this point I have no idea why anyone would put anything in us-east-1. Also isolation is not as good as they would have you believe: I am unable to login to AWS Quicksight in us-west-2...

Been in us-east-1 for a long time. Things like Direct Connect and other integrations aren't easy or cheap to move and when you have other, bigger priorities, moving regions is not an easy decision to prioritize.

Re: AWS us-east-1 outage

#254
post #13

I love that every time this happens, 100% of the services on https://status.aws.amazon.com are green.

It's like trying to get the truth out of a kid that caused some trouble.

Mom: Alexa, did you break something?

Alexa: No.

M: Really? What's this? 500 Internal server error

A: ok maybe management console is down

M: Anything else?

A: ...

A: ... ok maybe cloudwatch logs

M: Ah hah. What else?

A: That's it, I swear!

M: 503 ClientError

A: ...well okay secretsmanager might be busted too...

Re: AWS us-east-1 outage

#255

I'm now getting failures searching for products on Amazon.com itself. This is somewhat surprising, as the narrative always was that Amazon didn't do a great job of dogfooding their own cloud platform.

Update: I'm also getting Internal Errors trying to log into the Amazon.com site now as well.

Re: AWS us-east-1 outage

#256

It's funny that the first place I go to learn about the outage is Hacker News and not https://status.aws.amazon.com/ (it's still reports everything to be "operating normally"...)

I made sure our incident response plan includes checking Hacker News and Twitter for actual updates and information.

As of right now, this thread and one update from a twitter user, https://twitter.com/SiteRelEnby/status/1468253604876333059 are all we have. I went into disaster recovery mode when I saw our traffic dropped to 0 suddenly at 10:30am ET. That was just the SQS/something else preventing our ELB logs from being extracted to DataDog though.

Re: AWS us-east-1 outage

#257

The Amazon.com storefront was giving me issues loading search results — this is the worst possible time of year for Amazon to have issues. It’s horrifying and awesome to imagine hundreds of thousands (if not millions) of dollars of lost orders an hour — just from sluggish load times. Hugops to those dealing with this.

Third worst time. It's not BFCM and it's not the week before Christmas; from prior high-volume ecommerce experience I suspect their purchase rate is elevated at this time but nowhere near those two peaks.

Re: AWS us-east-1 outage

#258
post #245

Looks like they've acknowledged it on the status page now. https://status.aws.amazon.com/ > 8:22 AM PST We are investigating increased error rates for the AWS Management Console. > 8:26 AM PST We are experiencing API and console issues in the US-EAST-1 Region. We have identified root cause and we are actively working towards recovery. This issue is affecting the global console landing page, which is also hosted in US…

https://status.aws.amazon.com/ still shows all green for me

Re: AWS us-east-1 outage

#259
post #212

A former colleague told me years ago that us-east-1 is basically the guinea pig where changes get tested before being rolled out to the other regions, and as a result is less stable than the others. Does anyone know if there's any truth to this?

This is not true. Lambda updates us-east-1 last.
Post reply on HN