Live data from Hacker News

AWS's us-east-1 region is experiencing issues

health.aws.amazon.com

21–30 of 168 posts

Re: AWS's us-east-1 region is experiencing issues

#22

Based on our telemetry, this started as NXDOMAINs for sqs.us-east-1.amazonaws.com beginning in modest volumes at 20:43 UTC and becoming a total outage at 20:48 UTC. Naturally, it was completely resolved by 20:57, 5 minutes before anything was posted in the "Personal Health Dashboard" in the AWS console. It takes a while to find a Vice President, I guess.

Or perhaps triaging, root-causing, and fixing the issue is the highest-order bit?

Re: AWS's us-east-1 region is experiencing issues

#23
post #8

This is why you are strongly urged not to rely on one region or AZ.

Seems like it would be conflict of interest to increase robustness of single AZ (so it never goes down or has its own redundancy) vs. increased revenues from multi AZ deployment. What's the point of cloud if we have to manage robustness of their own infrastructure. I can understand if that's due to natural disasters and earthquakes, but the idea should be that a single AZ should never go down barring extraordinary ci…

They would simply charge for the privilege. An EC2 'always on' or whatever option that enabled your instance to live migrate between availability zones would be a nice and expensive option.

Re: AWS's us-east-1 region is experiencing issues

#24

Based on our telemetry, this started as NXDOMAINs for sqs.us-east-1.amazonaws.com beginning in modest volumes at 20:43 UTC and becoming a total outage at 20:48 UTC. Naturally, it was completely resolved by 20:57, 5 minutes before anything was posted in the "Personal Health Dashboard" in the AWS console. It takes a while to find a Vice President, I guess.

Or perhaps triaging, root-causing, and fixing the issue is the highest-order bit?

It definitely is. For an issue like this, you will see relevant teams and delegates looped in very quickly. Getting approved wording about an outage requires some very senior people though. Often they have to be paged in as well.

Having worked at a few other large tech companies now -- Amazon's incident response process is honestly great. It's one of the things I miss about working there.

Re: AWS's us-east-1 region is experiencing issues

#25

Based on our telemetry, this started as NXDOMAINs for sqs.us-east-1.amazonaws.com beginning in modest volumes at 20:43 UTC and becoming a total outage at 20:48 UTC. Naturally, it was completely resolved by 20:57, 5 minutes before anything was posted in the "Personal Health Dashboard" in the AWS console. It takes a while to find a Vice President, I guess.

Or perhaps triaging, root-causing, and fixing the issue is the highest-order bit?

sure, but if those people are updating the status pages to say something isn't right and we're looking into it, we're doomed.

Re: AWS's us-east-1 region is experiencing issues

#26

I can't help but wonder, with the increases in attrition across the industry, are we hitting some kind of tipping point where the institutional knowledge in these massive tech corporations is disappearing? Mistakes happen all the time but when all the people who intimately know how these systems work leave for other opportunities, disasters are bound to happen more and more.

Just like the tech priests in Warhammer 40k, keeping occult old engineering, thatno one could build anymore, running

Re: AWS's us-east-1 region is experiencing issues

#27

Based on our telemetry, this started as NXDOMAINs for sqs.us-east-1.amazonaws.com beginning in modest volumes at 20:43 UTC and becoming a total outage at 20:48 UTC. Naturally, it was completely resolved by 20:57, 5 minutes before anything was posted in the "Personal Health Dashboard" in the AWS console. It takes a while to find a Vice President, I guess.

Or perhaps triaging, root-causing, and fixing the issue is the highest-order bit?

Different people have different responsibilities. At Amazon scale, the comms and people doing a deep dive to fix stuff will not be the same.

Re: AWS's us-east-1 region is experiencing issues

#29

Earlier quoted context omitted.

Or perhaps triaging, root-causing, and fixing the issue is the highest-order bit?

It definitely is. For an issue like this, you will see relevant teams and delegates looped in very quickly. Getting approved wording about an outage requires some very senior people though. Often they have to be paged in as well. Having worked at a few other large tech companies now -- Amazon's incident response process is honestly great. It's one of the things I miss about working there.

This. We have a 4-person team and posted our own incident about this 7 minutes before Amazon did. Surely they can aim a little higher.

Re: AWS's us-east-1 region is experiencing issues

#30

What’s up with all of the multi-platform outages lately? Seems abnormal looking at historical data. Are there issues affecting the internet backbone or something? Or just a coincidence?

Important to keep in mind that AWS has 250 services in 84 Availability Zones in 26 regions.

This outage is reportedly impacting 5 services in 1 region.

For those impacted, pretty terrible. But as a heavy user of AWS, I’ve seen these notices posted multiple times on HN and haven’t been impacted by one yet.

Post reply on HN