Live data from Hacker News

Amazon EC2 currently down. Affecting Heroku, Reddit, Others

status.aws.amazon.com

191–200 of 311 posts

Re: Amazon EC2 currently down. Affecting Heroku, Reddit, Others

#193

Update from http://status.aws.amazon.com/ : 2:20 PM PDT We've now restored performance for about half of the volumes that experienced issues. Instances that were attached to these recovered volumes are recovering. We're continuing to work on restoring availability and performance for the volumes that are still degraded. We also want to add some detail around what customers using ELB may have experienced. Customers wi…

This is infuriating. Our instances aren't ELB-backed, so we have an entire AZ sitting idle while the rest of our instances are overwhelmed with traffic. Why did they make this decision for us?

I'm normally an Amazon apologist when outages happen, but this is absolutely ridiculous.

Re: Amazon EC2 currently down. Affecting Heroku, Reddit, Others

#194

Unfortunate. I have to deal with a number of folks who will be overjoyed to read this news when their tech cartel vendor of choice forwards it this evening. There's a huge contingent of currently endangered infrastructure folks (and vendors who feed off them) out there who throw a party every time AWS has a visible outage.

I'd question if they are truly wrong to send out news like this? Basically you have to weigh the realities of the 'cloud' with the perception.

The perception that many people has is that somehow the 'cloud' is a magical up-time device that will save you money in droves.

The reality is that for many companies with mid-level traffic and aren't a start-up with a billion users, the 'old' style tech might very well be your best option.

Re: Amazon EC2 currently down. Affecting Heroku, Reddit, Others

#195
post #31

The N. Virginia datacenter has been historically unreliable. I moved my personal projects to the West Coast (Oregon and N. California) and I have seen no significant issues in the past year. N. Virginia is both cheaper and closer to the center of mass of the developed world. I'm surprised Amazon hasn't managed to make it more reliable.

I don't understand why anyone's site is only in one datacenter. i thought the point of AWS was that it was distributed with fault tolerance? Why don't they distribute all the sites/apps across all their centers?

It takes development/engineering resources, and additional hardware resources to make your architecture more fault-tolerant and to maintain this fault-tolerance over long periods of time.

Weigh this against the estimated costs of your application going down occasionally. It's really only economical for the largest applications (Netflix, etc.) to build these systems.

Re: Amazon EC2 currently down. Affecting Heroku, Reddit, Others

#196
post #190

Earlier quoted context omitted.

I'm guessing your Chaos Gorilla helped to harden your architecture against this threat. Since you've mostly recovered, how did your system do? Are there side-cases that Chaos Gorilla didn't touch?

EDIT: I WAS WRONG. Chaos Monkey and Chaos Gorilla both exist and simulate different forms of chaos.

Monkey takes down single hosts; Gorilla simulates the loss of an entire AZ.

Re: Amazon EC2 currently down. Affecting Heroku, Reddit, Others

#197
post #190

Earlier quoted context omitted.

I'm guessing your Chaos Gorilla helped to harden your architecture against this threat. Since you've mostly recovered, how did your system do? Are there side-cases that Chaos Gorilla didn't touch?

EDIT: I WAS WRONG. Chaos Monkey and Chaos Gorilla both exist and simulate different forms of chaos.

Chaos Monkey takes down instances and such. Chaos Gorilla takes down entire AZs. :)

Re: Amazon EC2 currently down. Affecting Heroku, Reddit, Others

#198
post #190

Earlier quoted context omitted.

I'm guessing your Chaos Gorilla helped to harden your architecture against this threat. Since you've mostly recovered, how did your system do? Are there side-cases that Chaos Gorilla didn't touch?

EDIT: I WAS WRONG. Chaos Monkey and Chaos Gorilla both exist and simulate different forms of chaos.

"Create More Failures

Currently, Netflix uses a service called "Chaos Monkey" to simulate service failure. Basically, Chaos Monkey is a service that kills other services. We run this service because we want engineering teams to be used to a constant level of failure in the cloud. Services should automatically recover without any manual intervention. We don't however, simulate what happens when an entire AZ goes down and therefore we haven't engineered our systems to automatically deal with those sorts of failures. Internally we are having discussions about doing that and people are already starting to call this service "Chaos Gorilla"."

http://techblog.netflix.com/2011_04_01_archive.html

Re: Amazon EC2 currently down. Affecting Heroku, Reddit, Others

#199
post #190

Earlier quoted context omitted.

EDIT: I WAS WRONG. Chaos Monkey and Chaos Gorilla both exist and simulate different forms of chaos.

"Create More Failures Currently, Netflix uses a service called "Chaos Monkey" to simulate service failure. Basically, Chaos Monkey is a service that kills other services. We run this service because we want engineering teams to be used to a constant level of failure in the cloud. Services should automatically recover without any manual intervention. We don't however, simulate what happens when an entire AZ goes down…

My apologies! I was wrong.

Re: Amazon EC2 currently down. Affecting Heroku, Reddit, Others

#200
post #192

Convenient that we're too backward to use AWS. That means everyone can at least talk about it here when AWS is down.

I'm not familiar with HN's technical stack (other than arc), how has it scaled as the community grew over the years?
Post reply on HN