Summary of the Amazon Kinesis Event in the Northern Virginia (US-East-1) Region
1–10 of 153 posts
Re: Summary of the Amazon Kinesis Event in the Northern Virginia (US-East-1) Region
#2At 9:39 AM PST, we were able to confirm a root cause [...] the new capacity had caused all of the servers in the fleet to exceed the maximum number of threads allowed by an operating system configuration.
Re: Summary of the Amazon Kinesis Event in the Northern Virginia (US-East-1) Region
#3Poetry.
Then, to be fair:
> We have a back-up means of updating the Service Health Dashboard that has minimal service dependencies. While this worked as expected, we encountered several delays during the earlier part of the event in posting to the Service Health Dashboard with this tool, as it is a more manual and less familiar tool for our support operators. To ensure customers were getting timely updates, the support team used the Personal Health Dashboard to notify impacted customers if they were impacted by the service issues.
I'm curious if anyone here actually got one of these.
Re: Summary of the Amazon Kinesis Event in the Northern Virginia (US-East-1) Region
#4Re: Summary of the Amazon Kinesis Event in the Northern Virginia (US-East-1) Region
#5Re: Summary of the Amazon Kinesis Event in the Northern Virginia (US-East-1) Region
#6> During the early part of this event, we were unable to update the Service Health Dashboard because the tool we use to post these updates itself uses Cognito, which was impacted by this event. Poetry. Then, to be fair: > We have a back-up means of updating the Service Health Dashboard that has minimal service dependencies. While this worked as expected, we encountered several delays during the earlier part of the ev…
Re: Summary of the Amazon Kinesis Event in the Northern Virginia (US-East-1) Region
#7> During the early part of this event, we were unable to update the Service Health Dashboard because the tool we use to post these updates itself uses Cognito, which was impacted by this event. Poetry. Then, to be fair: > We have a back-up means of updating the Service Health Dashboard that has minimal service dependencies. While this worked as expected, we encountered several delays during the earlier part of the ev…
Re: Summary of the Amazon Kinesis Event in the Northern Virginia (US-East-1) Region
#8Re: Summary of the Amazon Kinesis Event in the Northern Virginia (US-East-1) Region
#9> During the early part of this event, we were unable to update the Service Health Dashboard because the tool we use to post these updates itself uses Cognito, which was impacted by this event. Poetry. Then, to be fair: > We have a back-up means of updating the Service Health Dashboard that has minimal service dependencies. While this worked as expected, we encountered several delays during the earlier part of the ev…
This won't be a first. The status page was hosted in S3. It is hilarious in the hindsight, but understandable.
Is it really? I get the value of eating your own dogfood, it improves things a lot.
But your status page? Such a high importance, low difficulty thing to build that dogfeeding it gives you small amount of benefits (dogfeed something bigger/more complex instead) in the good case, and high amount of drawback when things go wrong (like when your infrastructure goes down, so does your status page). So what's the point?
Re: Summary of the Amazon Kinesis Event in the Northern Virginia (US-East-1) Region
#10The one thing I want to know in cases like this is: why did it affect multiple Availability Zones? Making a resource multi-AZ is a significant additional cost (and often involves additional complexity) and we really need to be confident that typical observed outages would actually have been mitigated in return.