Live data from Hacker News

Amazon EC2 currently down. Affecting Heroku, Reddit, Others

status.aws.amazon.com

181–190 of 311 posts

Re: Amazon EC2 currently down. Affecting Heroku, Reddit, Others

#181
post #173

Earlier quoted context omitted.

Our multi-AZ RDS instance did not failover correctly this time although it has in the past..

same here, no fail-over

I had one multi-az failover correctly, however the security group was refusing connections to the web servers ec2 security group. I had to manually add in the private ips of the ec2 instances. It appears the API issue is affecting security group to ip lookups.

Re: Amazon EC2 currently down. Affecting Heroku, Reddit, Others

#182
Update from http://status.aws.amazon.com/:

2:20 PM PDT We've now restored performance for about half of the volumes that experienced issues. Instances that were attached to these recovered volumes are recovering. We're continuing to work on restoring availability and performance for the volumes that are still degraded.

We also want to add some detail around what customers using ELB may have experienced. Customers with ELBs running in only the affected Availability Zone may be experiencing elevated error rates and customers may not be able to create new ELBs in the affected Availability Zone. For customers with multi-AZ ELBs, traffic was shifted away from the affected Availability Zone early in this event and they should not be seeing impact at this time.

Re: Amazon EC2 currently down. Affecting Heroku, Reddit, Others

#183
post #143

Why the fuck do the two most critical services (ELB and Console) have depends on their historically most unreliable pile of shit (EBS)?

Seriously THIS has to be addressed.

I can tolerate EC2/EBS going down but why on earth is ELB/Console always going down at the same time ?

Re: Amazon EC2 currently down. Affecting Heroku, Reddit, Others

#185
post #31

The N. Virginia datacenter has been historically unreliable. I moved my personal projects to the West Coast (Oregon and N. California) and I have seen no significant issues in the past year. N. Virginia is both cheaper and closer to the center of mass of the developed world. I'm surprised Amazon hasn't managed to make it more reliable.

I don't understand why anyone's site is only in one datacenter. i thought the point of AWS was that it was distributed with fault tolerance? Why don't they distribute all the sites/apps across all their centers?

Re: Amazon EC2 currently down. Affecting Heroku, Reddit, Others

#186
post #91

Earlier quoted context omitted.

Yet netflix is currently down, so maybe there is a problem with the chimp army?

The problem is likely the same as usual: if the damn control plane is down, it doesn't matter how robust your failover architecture is, because your requests to bring up new machines go unanswered. There's pretty much no way to architect around that one as an AWS user (apart from going fully multi-cloud, but "nobody" actually does that, at least at scale), and I'm kind of shocked that those bits of AWS are still not…

Apparently iCloud is multi cloud (AWS and Azure).

But regardless it's not like all of EC2 went down just one or two AZs. So why couldn't traffic be migrated transprently to other AZs/regions ?

Re: Amazon EC2 currently down. Affecting Heroku, Reddit, Others

#187

This is why I trust Google's data centers http://www.google.com/about/datacenters/gallery/#/

HN really needs to hold a "dumbest comments of the year" award. This has to be a strong contender.

HN really needs to hold a "snarkiest response of the year" award.

Re: Amazon EC2 currently down. Affecting Heroku, Reddit, Others

#188
post #31

The N. Virginia datacenter has been historically unreliable. I moved my personal projects to the West Coast (Oregon and N. California) and I have seen no significant issues in the past year. N. Virginia is both cheaper and closer to the center of mass of the developed world. I'm surprised Amazon hasn't managed to make it more reliable.

Their CoLo space is the same space shared by AOL and a few other big name tech companies. It's right next to the Greenway, just before you reach IAD going northeast. That CoLo facility seems pretty unreliable in the scheme of things; Verizon and Amazon both took major downtime this summer when a pretty hefty storm rolled through VA[1], but AOL's dedicated datacenters in the same 10 mile radius all experienced no down…

To be fair, the entire region was decimated by that storm. I didn't have power for 5 days. Much of the area was out. There was a ton of physical damage. That's not excusing them, they should do better, but that storm was like nothing I've experienced living in the area for 20 years.

Re: Amazon EC2 currently down. Affecting Heroku, Reddit, Others

#189

Earlier quoted context omitted.

This kind of thinking is poisonous. I know it's in good fun and it's fun to look for connections in things, but it is actually preposterous to think that Amazon would purposely disrupt wide swaths of highly paying customers for much of a day to bury one story about bad customer relations. My guess is there are a lot of people working very hard to try and solve this problem right now, let's not belittle their efforts…

It was a joke. I made the same joke earlier today. No-one is seriously going to believe this.

HNers are extremely bad at getting jokes. See http://news.ycombinator.com/item?id=4677335

Re: Amazon EC2 currently down. Affecting Heroku, Reddit, Others

#190

Earlier quoted context omitted.

We (Netflix) have done a bunch of presentations on it which are on our slideshare page and across the internet. After this issue is over I can give a longer answer. In short, we've just evacuated the affected zone and are mostly recovered.

I'm guessing your Chaos Gorilla helped to harden your architecture against this threat. Since you've mostly recovered, how did your system do? Are there side-cases that Chaos Gorilla didn't touch?

EDIT: I WAS WRONG. Chaos Monkey and Chaos Gorilla both exist and simulate different forms of chaos.
Post reply on HN