Earlier quoted context omitted.
Our multi-AZ RDS instance did not failover correctly this time although it has in the past..
same here, no fail-over
Amazon EC2 currently down. Affecting Heroku, Reddit, Others
181–190 of 311 posts
Re: Amazon EC2 currently down. Affecting Heroku, Reddit, Others
#1822:20 PM PDT We've now restored performance for about half of the volumes that experienced issues. Instances that were attached to these recovered volumes are recovering. We're continuing to work on restoring availability and performance for the volumes that are still degraded.
We also want to add some detail around what customers using ELB may have experienced. Customers with ELBs running in only the affected Availability Zone may be experiencing elevated error rates and customers may not be able to create new ELBs in the affected Availability Zone. For customers with multi-AZ ELBs, traffic was shifted away from the affected Availability Zone early in this event and they should not be seeing impact at this time.
Re: Amazon EC2 currently down. Affecting Heroku, Reddit, Others
#183Why the fuck do the two most critical services (ELB and Console) have depends on their historically most unreliable pile of shit (EBS)?
I can tolerate EC2/EBS going down but why on earth is ELB/Console always going down at the same time ?
Re: Amazon EC2 currently down. Affecting Heroku, Reddit, Others
#184Re: Amazon EC2 currently down. Affecting Heroku, Reddit, Others
#185The N. Virginia datacenter has been historically unreliable. I moved my personal projects to the West Coast (Oregon and N. California) and I have seen no significant issues in the past year. N. Virginia is both cheaper and closer to the center of mass of the developed world. I'm surprised Amazon hasn't managed to make it more reliable.
Re: Amazon EC2 currently down. Affecting Heroku, Reddit, Others
#186Earlier quoted context omitted.
Yet netflix is currently down, so maybe there is a problem with the chimp army?
The problem is likely the same as usual: if the damn control plane is down, it doesn't matter how robust your failover architecture is, because your requests to bring up new machines go unanswered. There's pretty much no way to architect around that one as an AWS user (apart from going fully multi-cloud, but "nobody" actually does that, at least at scale), and I'm kind of shocked that those bits of AWS are still not…
But regardless it's not like all of EC2 went down just one or two AZs. So why couldn't traffic be migrated transprently to other AZs/regions ?
Re: Amazon EC2 currently down. Affecting Heroku, Reddit, Others
#187Re: Amazon EC2 currently down. Affecting Heroku, Reddit, Others
#188The N. Virginia datacenter has been historically unreliable. I moved my personal projects to the West Coast (Oregon and N. California) and I have seen no significant issues in the past year. N. Virginia is both cheaper and closer to the center of mass of the developed world. I'm surprised Amazon hasn't managed to make it more reliable.
Their CoLo space is the same space shared by AOL and a few other big name tech companies. It's right next to the Greenway, just before you reach IAD going northeast. That CoLo facility seems pretty unreliable in the scheme of things; Verizon and Amazon both took major downtime this summer when a pretty hefty storm rolled through VA[1], but AOL's dedicated datacenters in the same 10 mile radius all experienced no down…
Re: Amazon EC2 currently down. Affecting Heroku, Reddit, Others
#189Earlier quoted context omitted.
This kind of thinking is poisonous. I know it's in good fun and it's fun to look for connections in things, but it is actually preposterous to think that Amazon would purposely disrupt wide swaths of highly paying customers for much of a day to bury one story about bad customer relations. My guess is there are a lot of people working very hard to try and solve this problem right now, let's not belittle their efforts…
It was a joke. I made the same joke earlier today. No-one is seriously going to believe this.
Re: Amazon EC2 currently down. Affecting Heroku, Reddit, Others
#190Earlier quoted context omitted.
We (Netflix) have done a bunch of presentations on it which are on our slideshare page and across the internet. After this issue is over I can give a longer answer. In short, we've just evacuated the affected zone and are mostly recovered.
I'm guessing your Chaos Gorilla helped to harden your architecture against this threat. Since you've mostly recovered, how did your system do? Are there side-cases that Chaos Gorilla didn't touch?