"If you're having SLA problems I feel bad for you son
I got two 9 problems cuz of us-east-1"
21–30 of 168 posts
"If you're having SLA problems I feel bad for you son
I got two 9 problems cuz of us-east-1"
Based on our telemetry, this started as NXDOMAINs for sqs.us-east-1.amazonaws.com beginning in modest volumes at 20:43 UTC and becoming a total outage at 20:48 UTC. Naturally, it was completely resolved by 20:57, 5 minutes before anything was posted in the "Personal Health Dashboard" in the AWS console. It takes a while to find a Vice President, I guess.
This is why you are strongly urged not to rely on one region or AZ.
Seems like it would be conflict of interest to increase robustness of single AZ (so it never goes down or has its own redundancy) vs. increased revenues from multi AZ deployment. What's the point of cloud if we have to manage robustness of their own infrastructure. I can understand if that's due to natural disasters and earthquakes, but the idea should be that a single AZ should never go down barring extraordinary ci…
Based on our telemetry, this started as NXDOMAINs for sqs.us-east-1.amazonaws.com beginning in modest volumes at 20:43 UTC and becoming a total outage at 20:48 UTC. Naturally, it was completely resolved by 20:57, 5 minutes before anything was posted in the "Personal Health Dashboard" in the AWS console. It takes a while to find a Vice President, I guess.
Or perhaps triaging, root-causing, and fixing the issue is the highest-order bit?
Having worked at a few other large tech companies now -- Amazon's incident response process is honestly great. It's one of the things I miss about working there.
Based on our telemetry, this started as NXDOMAINs for sqs.us-east-1.amazonaws.com beginning in modest volumes at 20:43 UTC and becoming a total outage at 20:48 UTC. Naturally, it was completely resolved by 20:57, 5 minutes before anything was posted in the "Personal Health Dashboard" in the AWS console. It takes a while to find a Vice President, I guess.
Or perhaps triaging, root-causing, and fixing the issue is the highest-order bit?
I can't help but wonder, with the increases in attrition across the industry, are we hitting some kind of tipping point where the institutional knowledge in these massive tech corporations is disappearing? Mistakes happen all the time but when all the people who intimately know how these systems work leave for other opportunities, disasters are bound to happen more and more.
Based on our telemetry, this started as NXDOMAINs for sqs.us-east-1.amazonaws.com beginning in modest volumes at 20:43 UTC and becoming a total outage at 20:48 UTC. Naturally, it was completely resolved by 20:57, 5 minutes before anything was posted in the "Personal Health Dashboard" in the AWS console. It takes a while to find a Vice President, I guess.
Or perhaps triaging, root-causing, and fixing the issue is the highest-order bit?
Earlier quoted context omitted.
Or perhaps triaging, root-causing, and fixing the issue is the highest-order bit?
It definitely is. For an issue like this, you will see relevant teams and delegates looped in very quickly. Getting approved wording about an outage requires some very senior people though. Often they have to be paged in as well. Having worked at a few other large tech companies now -- Amazon's incident response process is honestly great. It's one of the things I miss about working there.
What’s up with all of the multi-platform outages lately? Seems abnormal looking at historical data. Are there issues affecting the internet backbone or something? Or just a coincidence?
This outage is reportedly impacting 5 services in 1 region.
For those impacted, pretty terrible. But as a heavy user of AWS, I’ve seen these notices posted multiple times on HN and haven’t been impacted by one yet.