Live data from Hacker News

Amazon AWS had a power failure, their backup generators failed

twitter.com

81–90 of 106 posts

Re: Amazon AWS had a power failure, their backup generators failed

#81

This seems to be getting slightly overblown in that thread. To be clear, this impacted one datacenter out of ten that make up one availability zone out of six in AWS’s us-east-1 region. So we are talking 2-3% at most of that region’s capacity was impacted. I haven’t seen a report yet on exactly why their generator failed, but from what I’ve heard, the power failed, and the backup generator kicked in and ran fine for…

Don't disagree with your point that this is overblown, but here's an important related point

"Your nines are not my nines" - https://rachelbythebay.com/w/2019/07/15/giant/

Re: Amazon AWS had a power failure, their backup generators failed

#82

This seems to be getting slightly overblown in that thread. To be clear, this impacted one datacenter out of ten that make up one availability zone out of six in AWS’s us-east-1 region. So we are talking 2-3% at most of that region’s capacity was impacted. I haven’t seen a report yet on exactly why their generator failed, but from what I’ve heard, the power failed, and the backup generator kicked in and ran fine for…

agree. like you imply, too many people rely heavily on their providers for business critical things like backups and redundancy. while generally big providers to a good job at this (and aws certainly does a good job at this), it does not mean there is a guarantee of any kind failures won't ever occur. thus the need to heavily invest in failure resistant technologies upon this borrowed infrastructure is arguably more…

I suspect the four days was how long it took to confirm that their particular EBS volume was definitely not recoverable. Soon after the outage ended on Saturday morning it was clear that some EBS volumes were not recovering quickly, and support's advice at that time was to rebuild/recover if you needed to still be online.

When a datacenter loses power like this, a few of the storage arrays will just not come back online. But another few will take time to run through their corruption recovery process, and it may take a long time and some service by a human (eg parts replacement, etc) before they can be certain a particular volume is not recoverable. Given their scale and the timing, at the beginning of a holiday weekend, four days is annoying, but not bad.

Re: Amazon AWS had a power failure, their backup generators failed

#83

Earlier quoted context omitted.

agree. like you imply, too many people rely heavily on their providers for business critical things like backups and redundancy. while generally big providers to a good job at this (and aws certainly does a good job at this), it does not mean there is a guarantee of any kind failures won't ever occur. thus the need to heavily invest in failure resistant technologies upon this borrowed infrastructure is arguably more…

But the reason I pay AWS is so that I don't have to hire a team to take care of backups and redundancy on my side. If they can't be relied on, a lot of the justification for their cost markup goes out the window.

It sounds like you might misunderstand the product you are buying from them. They are very clear about the reliability of EBS (1 in 1000 volumes will fail during a year of uptime), and they provide a really easy way to back things up, and there are tools available to schedule automated backup rotation. So I'm not sure what more you expect. AWS can't possibly know what your needs are for backup and restore for a particular EBS volume. If you want data durability, use S3.

Re: Amazon AWS had a power failure, their backup generators failed

#84

This seems to be getting slightly overblown in that thread. To be clear, this impacted one datacenter out of ten that make up one availability zone out of six in AWS’s us-east-1 region. So we are talking 2-3% at most of that region’s capacity was impacted. I haven’t seen a report yet on exactly why their generator failed, but from what I’ve heard, the power failed, and the backup generator kicked in and ran fine for…

> This seems to be getting slightly overblown in that thread. To be clear, this impacted one datacenter out of ten that make up one availability zone out of six in AWS’s us-east-1 region. So we are talking 2-3% at most of that region’s capacity was impacted.

For how many users was this 100% of their business?

Single-digit outage percentages for cloud services like AWS look like no big deal from the big perspective, which means that they often aren't a high priority. But when you're one of the customers in the 2-3%, it is a big deal, and you want it to be a high priority.

Small hosting providers may have only 3 9s of availability, but when your site goes down their world stops until it's fixed. I've seen reports of 5 9s of availability from Amazon, but when your site goes down, their alerts are still all green and they'll call you back at their leisure.

Re: Amazon AWS had a power failure, their backup generators failed

#85
People tend to make a habit of treating AWS like a VPS provider, in my experience. And you can skate by on this for a while, but it really isn't designed for that. And that will, eventually, lead to pain and suffering.

Sometimes instances just go out to lunch. Sometimes an AZ goes down. Chaos Monkey isn't just a good idea, it's required for reliable operation.

But please, please, if you are going to treat AWS like a VPS, at least don't do it in us-east-1! It seems to have more outages.

Our setup related to instances and EBS includes: At least 2 instances in different AZs, an ELB in front of them, a backup running at our hosting facility (though this could just be a different AWS zone, or different provider), and DNS with full-paper-path health checks that switch DNS over to the colo servers if any component of the primary fails.

Re: Amazon AWS had a power failure, their backup generators failed

#86

This seems to be getting slightly overblown in that thread. To be clear, this impacted one datacenter out of ten that make up one availability zone out of six in AWS’s us-east-1 region. So we are talking 2-3% at most of that region’s capacity was impacted. I haven’t seen a report yet on exactly why their generator failed, but from what I’ve heard, the power failed, and the backup generator kicked in and ran fine for…

Don't disagree with your point that this is overblown, but here's an important related point "Your nines are not my nines" - https://rachelbythebay.com/w/2019/07/15/giant/

If I recall correctly this is nicely covered in the book "Release it!" which BTW I recommend.

Re: Amazon AWS had a power failure, their backup generators failed

#87

We had a similar problem at our hosting facility last winter during that "Ice Vortex" storm that was all over the news. The facilities guys had been very proud of never having a power outage, I've had servers with them for 15 years now. The morning of the worst of the storm, we completely lose access to all services at that facility. Super unusual, everything we have there is redundant. So I do some minor investigati…

Oh, I forgot to mention that their solution was to run off generators for the impacted UPSes until repairs could be completed. I think one of them ended up running for 2 weeks. It seems strange that replacement parts for this UPS were so hard to source. That's why it was down in the first place. Someone making them by hand?

Re: Amazon AWS had a power failure, their backup generators failed

#88

This seems to be getting slightly overblown in that thread. To be clear, this impacted one datacenter out of ten that make up one availability zone out of six in AWS’s us-east-1 region. So we are talking 2-3% at most of that region’s capacity was impacted. I haven’t seen a report yet on exactly why their generator failed, but from what I’ve heard, the power failed, and the backup generator kicked in and ran fine for…

Some AWS services can't span across multiple availability zones like EMR.

Even in these cases, however, there are usually architectural methods for mitigating this kind of failure. In the case of EMR, using transient clusters with state stored in durable file stores such as EMRFS and vanilla S3 is one such option.

It is inevitable that systems will fail. The best the industry can do is work to reduce the number of failures and understand failure modes well so that they can be planned for. AWS does a very good job of this in my experience.

Regardless of whether applications are hosted in the cloud, on premises, co-located or in some hybrid configuration, it's important to design for that inevitable failure and keep business decision makers in the loop while doing so. Understanding requirements around RPO and RTO are extremely important in developing an architecture which meets the needs of the business, yet is still cost effective.

Re: Amazon AWS had a power failure, their backup generators failed

#89

Earlier quoted context omitted.

agree. like you imply, too many people rely heavily on their providers for business critical things like backups and redundancy. while generally big providers to a good job at this (and aws certainly does a good job at this), it does not mean there is a guarantee of any kind failures won't ever occur. thus the need to heavily invest in failure resistant technologies upon this borrowed infrastructure is arguably more…

But the reason I pay AWS is so that I don't have to hire a team to take care of backups and redundancy on my side. If they can't be relied on, a lot of the justification for their cost markup goes out the window.

If you replace AWS with Heroku in your statement, I agree. Heroku abstracts the redundant AWS resources for you so you can just “run your app”. However, Heroku also had a huge outage. That is way more problematic as far as I am concerned.

Re: Amazon AWS had a power failure, their backup generators failed

#90

We had a similar problem at our hosting facility last winter during that "Ice Vortex" storm that was all over the news. The facilities guys had been very proud of never having a power outage, I've had servers with them for 15 years now. The morning of the worst of the storm, we completely lose access to all services at that facility. Super unusual, everything we have there is redundant. So I do some minor investigati…

I had a similar situation in the early 2000's at NTT/Verio in Sterling, VA. They lost power because someone dug through a utility line and introduced a ground a fault into their system. They switched to generators and then against protocol (basically human error) tried to force override a transfer switch onto their other utility which ended up killing the generators. Eventually the UPSs were all drained, but we had t…

Power seems like the big wildcard in data center management. So tough to properly test your failover preparations, and so many different ways things can go wrong.

I know of a large company that had the data center emergency cutoff button next to the automatic doors on the way out. Sure enough, a contractor hit it one day thinking it was the way to open the doors.

Post reply on HN