Live data from Hacker News

Summary of the Amazon DynamoDB Service Disruption

aws.amazon.com

71–80 of 89 posts

Re: Summary of the Amazon DynamoDB Service Disruption

#71

"It was impacting a relatively small number of customers, but we should have posted the green-i to the dashboard sooner than we did on Monday." Amazon. No. Please listen. You should never post a green-i. Green-i means nothing to anyone. It's a minimization of a problem. You should change the indicator to show that there is a problem. If there is a problem, that is what the indicator is for. Customers don't care about…

What's with this all-or-nothing attitude? What is so wrong about different severity levels? Green-i: Some small percentage of customers are affected, but most of you have nothing to worry about. So if you are a customer having problems and see this notice, you are re-assured that you are not crazy. And if you are NOT seeing problems, you probably won't. Amazon has not always been the best at posting this information…

Yep, especially yellow is labeled clearly as "Performance issues". It is basically lying to make things appear better than they really are.

Re: Summary of the Amazon DynamoDB Service Disruption

#72

"It was impacting a relatively small number of customers, but we should have posted the green-i to the dashboard sooner than we did on Monday." Amazon. No. Please listen. You should never post a green-i. Green-i means nothing to anyone. It's a minimization of a problem. You should change the indicator to show that there is a problem. If there is a problem, that is what the indicator is for. Customers don't care about…

What's with this all-or-nothing attitude? What is so wrong about different severity levels? Green-i: Some small percentage of customers are affected, but most of you have nothing to worry about. So if you are a customer having problems and see this notice, you are re-assured that you are not crazy. And if you are NOT seeing problems, you probably won't. Amazon has not always been the best at posting this information…

> Some small percentage of customers are affected, but most of you have nothing to worry about.

That's absolutely not what it means practically for any of the recent outages. It has come to mean "the service is down but not hard enough for us to admit it".

Re: Summary of the Amazon DynamoDB Service Disruption

#73
post #3

We did not have detailed enough monitoring for this dimension (membership size), and didn’t have enough capacity allocated to the metadata service to handle these much heavier requests. As much as I admire and rely on AWS' scale to build architectures and fault tolerant applications, it can't be ignored that the marketing towards going "full cloud" doesn't take into account how hard it is to build resilient architect…

How is it any harder to do in the cloud than on a rack in a warehouse? At least you don't have to muck about with cables and phoning power companies up.

You call up IBM. You ask for a mainframe solution for two sites. You get experts to set it up for you with your application and such. You don't worry about downtime again for at least 30 years.

You call up Bull, Fujitsu, or Unisys for the same thing.

You call up HP. You ask for a NonStop solution. You get same thing for at least 20 years.

You call up VMS Software. You ask for an OpenVMS cluster. You get same thing for at least 17 years.

Well-designed OS's, software, and hardware did cloud-style stuff for a long time before cloud existed without the downtime. Cloud certainly brought price down and flexibility up. Yet, these clouds haven't matched 70-80's technology in uptime yet despite all the brains and money thrown at them. That's a fact.

So, shouldn't be used for anything mission critical where downtime costs lots of money.

Re: Summary of the Amazon DynamoDB Service Disruption

#74

"It was impacting a relatively small number of customers, but we should have posted the green-i to the dashboard sooner than we did on Monday." Amazon. No. Please listen. You should never post a green-i. Green-i means nothing to anyone. It's a minimization of a problem. You should change the indicator to show that there is a problem. If there is a problem, that is what the indicator is for. Customers don't care about…

What's with this all-or-nothing attitude? What is so wrong about different severity levels? Green-i: Some small percentage of customers are affected, but most of you have nothing to worry about. So if you are a customer having problems and see this notice, you are re-assured that you are not crazy. And if you are NOT seeing problems, you probably won't. Amazon has not always been the best at posting this information…

> By 2:37am PDT, the error rate in customer requests to DynamoDB had risen far beyond any level experienced in the last 3 years, finally stabilizing at approximately 55%.

Yeah, a small green-i should cover it.

Re: Summary of the Amazon DynamoDB Service Disruption

#75

> but we should have posted the green-i to the dashboard sooner than we did Use full orange or even half orange circle or something else if you want to convey 'subset of customers may have issues'. If possible move 'having issue' items to the top - this way people do not have to scroll down to see whats broken.

Not enough granularity with a yellow icon. I feel like if only <1% of customers are affected, a yellow circle isn't warranted - that's for the 5-20% impact case

Unless you're one of the The whole point of status page is to to help determine if the issue you're experiencing is on your side or the SaaS, and not to show as much green as possible.

The way currently AWS status page works it simply fails to provide any functionality and might as well be shut down.

The colors on it start to change when you already see tons of articles about AWS being down.

Re: Summary of the Amazon DynamoDB Service Disruption

#76

Earlier quoted context omitted.

What's with this all-or-nothing attitude? What is so wrong about different severity levels? Green-i: Some small percentage of customers are affected, but most of you have nothing to worry about. So if you are a customer having problems and see this notice, you are re-assured that you are not crazy. And if you are NOT seeing problems, you probably won't. Amazon has not always been the best at posting this information…

Green traditionally means we're good, ready to go, etc. Non-intuitive to use any green indicator for a situation deserving a write-up like this. It's yellow at best.

During the actual outage this write up was about, didn't it go red?

Re: Summary of the Amazon DynamoDB Service Disruption

#77

Earlier quoted context omitted.

It's really easy to miss the indicator "i". I've been doing this for 4 years and it still doesn't catch my eye when scanning the page.

I agree 100%. That column of uniform green makes it easy to miss. It is a poor UI choice but can be easily fixed.

However it is hard to miss the difference in the column with details and the more link even upon a cursory scan.

Re: Summary of the Amazon DynamoDB Service Disruption

#78

Earlier quoted context omitted.

How is it any harder to do in the cloud than on a rack in a warehouse? At least you don't have to muck about with cables and phoning power companies up.

You call up IBM. You ask for a mainframe solution for two sites. You get experts to set it up for you with your application and such. You don't worry about downtime again for at least 30 years. You call up Bull, Fujitsu, or Unisys for the same thing. You call up HP. You ask for a NonStop solution. You get same thing for at least 20 years. You call up VMS Software. You ask for an OpenVMS cluster. You get same thing fo…

I've seen NonStop solution failing due to completely mundane reason of insufficient disk space after a burst of transactions. One condition for those 30-year uptimes is also a 100% predictable environment.

Re: Summary of the Amazon DynamoDB Service Disruption

#79

"It was impacting a relatively small number of customers, but we should have posted the green-i to the dashboard sooner than we did on Monday." Amazon. No. Please listen. You should never post a green-i. Green-i means nothing to anyone. It's a minimization of a problem. You should change the indicator to show that there is a problem. If there is a problem, that is what the indicator is for. Customers don't care about…

I'm starting to wonder how much more I should trust AWS. Is everything swept under the carpet?

e.g., not even attributed or thanks for root escalation vulnerabilities that I surfaced for them: https://news.ycombinator.com/item?id=10261507

Re: Summary of the Amazon DynamoDB Service Disruption

#80

Earlier quoted context omitted.

You call up IBM. You ask for a mainframe solution for two sites. You get experts to set it up for you with your application and such. You don't worry about downtime again for at least 30 years. You call up Bull, Fujitsu, or Unisys for the same thing. You call up HP. You ask for a NonStop solution. You get same thing for at least 20 years. You call up VMS Software. You ask for an OpenVMS cluster. You get same thing fo…

I've seen NonStop solution failing due to completely mundane reason of insufficient disk space after a burst of transactions. One condition for those 30-year uptimes is also a 100% predictable environment.

Never heard of that one. Funny stuff. Shows even top tier can be improved.

Edit to add: mainframes also run user and server type workloads. Some of those are predictable, some aren't. Bull's virtualize whole desktops. The mainframe as a whole, esp important services, are usually still available despite issues with those. For instance, my company splits stuff between critical on mainframe or AS/400's plus non-critical on whatever is useful ("best-of-breed" they say...). The critical stuff is either on the IBM stuff or leverages it in client-server setup. Those apps either always work or (rarely) they fail-safe in an obvious way that does no damage. Nobody I work with can remember those systems going down over 10 years they worked there. The other stuff regularly has issues across the board. The key difference is effective architecture and how it's implemented.

Post reply on HN