6:33 AM PDT We are investigating increased error rates for API requests in the US-EAST-1 Region.
Source: http://status.aws.amazon.com/
41–50 of 89 posts
6:33 AM PDT We are investigating increased error rates for API requests in the US-EAST-1 Region.
Source: http://status.aws.amazon.com/
Hmm, we are seeing very high request latencies with DynamoDB in multiple accounts and can't even load the DynamoDB console...Anyone else?
As this post is on, Dynamo is experiencing error again: 6:33 AM PDT We are investigating increased error rates for API requests in the US-EAST-1 Region. Source: http://status.aws.amazon.com/
Does anyone keep stats on service outages for AWS, Axure, Google, et al?
Most of the historical data isn't shown publicly but I could probably do something with that.
Unavailable servers continued to retry requests for membership data, maintaining high load on the metadata service. It sounds like the retries exacerbated the situation. I wonder if they use exponential backoff?
> I wonder if they use exponential backoff? More importantly, did they use randomised exponential backoff? Having all the retires hitting at the same time can lead to a pulses of outages until things settle down.
I'd say the Dynamo team are well aware of what they should have been doing, and kicking themselves for not foreseeing this cascading-failure case. ouch!
A pretty good post mortem, but one worrisome final comment: "For us, availability is the most important feature of DynamoDB" I would think that durability should be the most important feature of DynamoDB; better to have intermittent outage, or reduced capacity, but not have data loss - but perhaps there is something about DynamoDB which suggests availability is rated more highly than durability?
I suspect they're simply trying to reassure customers that this size of outage won't happen again, not make some kind of deep mission statement..
Amazon. No. Please listen. You should never post a green-i. Green-i means nothing to anyone. It's a minimization of a problem. You should change the indicator to show that there is a problem. If there is a problem, that is what the indicator is for. Customers don't care about whatever weird internal political pressure causes you to want to show as little information as possible on that dashboard. Customers want to know that there is a problem, and we will totally be able to figure out if it is small or large on our own.
"It was impacting a relatively small number of customers, but we should have posted the green-i to the dashboard sooner than we did on Monday." Amazon. No. Please listen. You should never post a green-i. Green-i means nothing to anyone. It's a minimization of a problem. You should change the indicator to show that there is a problem. If there is a problem, that is what the indicator is for. Customers don't care about…
Green-i: Some small percentage of customers are affected, but most of you have nothing to worry about. So if you are a customer having problems and see this notice, you are re-assured that you are not crazy. And if you are NOT seeing problems, you probably won't. Amazon has not always been the best at posting this information quickly enough, hence a commitment to be better at messaging for those customers
Yellow: A significant amount of customers is affected. Potentially serious problems, and if you haven't seen them yet, don't be surprised if you will. Start monitoring your service health and preparing for a failover to another region.
Red: FUBAR. Happens rarely, but when it does you'll probably know even before your application alarms - you'll notice when half the internet shuts down.
> but we should have posted the green-i to the dashboard sooner than we did Use full orange or even half orange circle or something else if you want to convey 'subset of customers may have issues'. If possible move 'having issue' items to the top - this way people do not have to scroll down to see whats broken.
"It was impacting a relatively small number of customers, but we should have posted the green-i to the dashboard sooner than we did on Monday." Amazon. No. Please listen. You should never post a green-i. Green-i means nothing to anyone. It's a minimization of a problem. You should change the indicator to show that there is a problem. If there is a problem, that is what the indicator is for. Customers don't care about…
What's with this all-or-nothing attitude? What is so wrong about different severity levels? Green-i: Some small percentage of customers are affected, but most of you have nothing to worry about. So if you are a customer having problems and see this notice, you are re-assured that you are not crazy. And if you are NOT seeing problems, you probably won't. Amazon has not always been the best at posting this information…
Customers are totally not interested in the fact that some other random customer may not be experiencing problems and that amazon has a wide variety of services, some of which may not be affected in certain geographies. Customers go to that panel with one express goal: to discover if the problems they are seeing with their servers could be related to them. It is a triage check. The question is not, is Amazon super great and are the availability engineers the bestest on average over time? The question is, what the fuck is going on?
Amazon has suffered catastrophic unavailability, and the green-i has appeared belatedly, an hour or more into the problem. Because as revealed in this post, it is manually put up by engineers. Manually!
Manually!