Is it because operating system configuration is managed by a different team within the organization?
Summary of the Amazon Kinesis Event in the Northern Virginia (US-East-1) Region
11–20 of 153 posts
Re: Summary of the Amazon Kinesis Event in the Northern Virginia (US-East-1) Region
#12> the new capacity had caused all of the servers in the fleet to exceed the maximum number of threads allowed by an operating system configuration. [...] We didn’t want to increase the operating system limit without further testing Is it because operating system configuration is managed by a different team within the organization?
Re: Summary of the Amazon Kinesis Event in the Northern Virginia (US-East-1) Region
#13> the new capacity had caused all of the servers in the fleet to exceed the maximum number of threads allowed by an operating system configuration. [...] We didn’t want to increase the operating system limit without further testing Is it because operating system configuration is managed by a different team within the organization?
Re: Summary of the Amazon Kinesis Event in the Northern Virginia (US-East-1) Region
#14First of all, we want to apologize for the impact this event caused for our customers. While we are proud of our long track record of availability with Amazon Kinesis, we know how critical this service is to our customers, their applications and end users, and their businesses. We will do everything we can to learn from this event and use it to improve our availability even further.
Then move on to explain...
Re: Summary of the Amazon Kinesis Event in the Northern Virginia (US-East-1) Region
#15Earlier quoted context omitted.
This won't be a first. The status page was hosted in S3. It is hilarious in the hindsight, but understandable.
> but understandable Is it really? I get the value of eating your own dogfood, it improves things a lot. But your status page? Such a high importance, low difficulty thing to build that dogfeeding it gives you small amount of benefits (dogfeed something bigger/more complex instead) in the good case, and high amount of drawback when things go wrong (like when your infrastructure goes down, so does your status page). S…
Re: Summary of the Amazon Kinesis Event in the Northern Virginia (US-East-1) Region
#16The one thing I want to know in cases like this is: why did it affect multiple Availability Zones? Making a resource multi-AZ is a significant additional cost (and often involves additional complexity) and we really need to be confident that typical observed outages would actually have been mitigated in return.
Indeed. We're paying (and designing our systems to work on multiple AZs) to reduce the risk of outages, but then their back-end services are reliant on services in a sole region?
I, as many people have, discovered this when something broke in one of the golden regions. In my case cloudfront and ACM.
Realistically you can’t trust one provider at all if you have high availability requirements.
The justification is apparently that the cloud is taking all this responsibility away from people but from personal experience running two cages of kit at two datacenters the TCO was lower and the reliability and availability higher. Possibly the largest cost is navigating Harry-Potter-esque pricing and automation laws. The only gain is scaling past those two cages.
Edit: I should point out however that an advantage of the cloud is actually being able to click a couple of buttons and get rid of two cages worth of DC equipment instantly if your product or idea doesn't work out!
Re: Summary of the Amazon Kinesis Event in the Northern Virginia (US-East-1) Region
#17> the new capacity had caused all of the servers in the fleet to exceed the maximum number of threads allowed by an operating system configuration. [...] We didn’t want to increase the operating system limit without further testing Is it because operating system configuration is managed by a different team within the organization?
Re: Summary of the Amazon Kinesis Event in the Northern Virginia (US-East-1) Region
#18I would have started the response with: First of all, we want to apologize for the impact this event caused for our customers. While we are proud of our long track record of availability with Amazon Kinesis, we know how critical this service is to our customers, their applications and end users, and their businesses. We will do everything we can to learn from this event and use it to improve our availability even fur…
Re: Summary of the Amazon Kinesis Event in the Northern Virginia (US-East-1) Region
#19> During the early part of this event, we were unable to update the Service Health Dashboard because the tool we use to post these updates itself uses Cognito, which was impacted by this event. Poetry. Then, to be fair: > We have a back-up means of updating the Service Health Dashboard that has minimal service dependencies. While this worked as expected, we encountered several delays during the earlier part of the ev…