Live data from Hacker News

Amazon AWS had a power failure, their backup generators failed

twitter.com

91–100 of 106 posts

Re: Amazon AWS had a power failure, their backup generators failed

#91

So dude is mad because he didn’t have a redundancy plan? You can take snapshots of EBS volumes which backs everything up to S3. They even tell you that EBS volumes can fail in the documentation. But blaming someone else is easier I guess...

How do you do restore in those scenarios? You have a whole mess of snapshots of data that's in various stages of being wrong.

I'd rather just put up with the agony of RDS to get a point in time restore and treat my instances' data as volatile.

Re: Amazon AWS had a power failure, their backup generators failed

#92
post #2

I wonder if they perform routine tests on their support infra: power, cooling, et al

Someone is always signing off on routine tests.

Same deal as those 50 point inspections your mechanic does: some things are easier to inspect than others, some people do a more thorough job than others, etc.

Re: Amazon AWS had a power failure, their backup generators failed

#93

This seems to be getting slightly overblown in that thread. To be clear, this impacted one datacenter out of ten that make up one availability zone out of six in AWS’s us-east-1 region. So we are talking 2-3% at most of that region’s capacity was impacted. I haven’t seen a report yet on exactly why their generator failed, but from what I’ve heard, the power failed, and the backup generator kicked in and ran fine for…

> This seems to be getting slightly overblown in that thread. To be clear, this impacted one datacenter out of ten that make up one availability zone out of six in AWS’s us-east-1 region. So we are talking 2-3% at most of that region’s capacity was impacted. For how many users was this 100% of their business? Single-digit outage percentages for cloud services like AWS look like no big deal from the big perspective, w…

Honestly, if you are running on a single AZ in aws you are setting yourself up for failure..

Re: Amazon AWS had a power failure, their backup generators failed

#94

Earlier quoted context omitted.

agree. like you imply, too many people rely heavily on their providers for business critical things like backups and redundancy. while generally big providers to a good job at this (and aws certainly does a good job at this), it does not mean there is a guarantee of any kind failures won't ever occur. thus the need to heavily invest in failure resistant technologies upon this borrowed infrastructure is arguably more…

But the reason I pay AWS is so that I don't have to hire a team to take care of backups and redundancy on my side. If they can't be relied on, a lot of the justification for their cost markup goes out the window.

AWS is only selling you infrastructure as a service, not a turnkey solution. It's up to you to combine and coordinate these services into a solution that delivers the capabilities (including backup, recovery and fault tolerance) appropriate to your needs. So while you don't need to hire a team to take care of backups and redundancy, you do need to provision and configure what is required so that their team can.

Re: Amazon AWS had a power failure, their backup generators failed

#95

This seems to be getting slightly overblown in that thread. To be clear, this impacted one datacenter out of ten that make up one availability zone out of six in AWS’s us-east-1 region. So we are talking 2-3% at most of that region’s capacity was impacted. I haven’t seen a report yet on exactly why their generator failed, but from what I’ve heard, the power failed, and the backup generator kicked in and ran fine for…

> This seems to be getting slightly overblown in that thread. To be clear, this impacted one datacenter out of ten that make up one availability zone out of six in AWS’s us-east-1 region. So we are talking 2-3% at most of that region’s capacity was impacted. For how many users was this 100% of their business? Single-digit outage percentages for cloud services like AWS look like no big deal from the big perspective, w…

> For how many users was this 100% of their business?

… but they still chose to ignore all of the prominent warnings and architectural guidance, not to mention avoiding use of the services which have HA built-in. I mean, I'm sympathetic to anyone who had a bad day with a forced learning experience but it's not like this is some dark secret.

Re: Amazon AWS had a power failure, their backup generators failed

#96

This seems to be getting slightly overblown in that thread. To be clear, this impacted one datacenter out of ten that make up one availability zone out of six in AWS’s us-east-1 region. So we are talking 2-3% at most of that region’s capacity was impacted. I haven’t seen a report yet on exactly why their generator failed, but from what I’ve heard, the power failed, and the backup generator kicked in and ran fine for…

Don't disagree with your point that this is overblown, but here's an important related point "Your nines are not my nines" - https://rachelbythebay.com/w/2019/07/15/giant/

But for a large number of 9s of people, AWS's 9s are their 9s.

Re: Amazon AWS had a power failure, their backup generators failed

#97
post #95

Earlier quoted context omitted.

> This seems to be getting slightly overblown in that thread. To be clear, this impacted one datacenter out of ten that make up one availability zone out of six in AWS’s us-east-1 region. So we are talking 2-3% at most of that region’s capacity was impacted. For how many users was this 100% of their business? Single-digit outage percentages for cloud services like AWS look like no big deal from the big perspective, w…

> For how many users was this 100% of their business? … but they still chose to ignore all of the prominent warnings and architectural guidance, not to mention avoiding use of the services which have HA built-in. I mean, I'm sympathetic to anyone who had a bad day with a forced learning experience but it's not like this is some dark secret.

> but they still chose to ignore all of the prominent warnings and architectural guidance

So that's it, blame the user and caveat emptor?

> not to mention avoiding use of the services which have HA built-in.

Many (all?) of these services tightly couple you to Amazon, so avoiding them is a very reasonable decision.

Re: Amazon AWS had a power failure, their backup generators failed

#98
post #95

Earlier quoted context omitted.

> For how many users was this 100% of their business? … but they still chose to ignore all of the prominent warnings and architectural guidance, not to mention avoiding use of the services which have HA built-in. I mean, I'm sympathetic to anyone who had a bad day with a forced learning experience but it's not like this is some dark secret.

> but they still chose to ignore all of the prominent warnings and architectural guidance So that's it, blame the user and caveat emptor? > not to mention avoiding use of the services which have HA built-in. Many (all?) of these services tightly couple you to Amazon, so avoiding them is a very reasonable decision.

> So that's it, blame the user and caveat emptor?

If I sell you a loaf of bread and you complain that it's not a sandwich, is it anything else?

> Many (all?) of these services tightly couple you to Amazon, so avoiding them is a very reasonable decision.

That's just a cop-out: checking the “multi-AZ” box in RDS completely avoided this problem with zero lock-in. If you're deploying containers, you have multiple options which are portable and avoid this completely. If you're deploying EC2 instances, again you have options with very limited lock-in (e.g. auto-scaling with multiple AZs).

More importantly, that's also a business decision: if you're that worried about lock-in you are accepting responsibility to operate the alternatives. For example, following industry-standard practice might suggest that you run everything in Kubernetes in multiple AZs, regions, or providers but it would never support running everything in a single AZ.

Re: Amazon AWS had a power failure, their backup generators failed

#99
Forgive me, What may be a stupid question.

As far as I can tell, the number of a Mechanical / Hardware Failure is far higher than Software. And it is always Power, UPS, Generator, BBU, Raid Card Failure etc.

Why is it that we keep hearing failure in these segment? And it doesn't seems anything have been done? Are there any Innovation happening in this space?

Re: Amazon AWS had a power failure, their backup generators failed

#100

This seems to be getting slightly overblown in that thread. To be clear, this impacted one datacenter out of ten that make up one availability zone out of six in AWS’s us-east-1 region. So we are talking 2-3% at most of that region’s capacity was impacted. I haven’t seen a report yet on exactly why their generator failed, but from what I’ve heard, the power failed, and the backup generator kicked in and ran fine for…

Don't disagree with your point that this is overblown, but here's an important related point "Your nines are not my nines" - https://rachelbythebay.com/w/2019/07/15/giant/

Reminds me of the hover text on this xkcd: https://xkcd.com/325/
Post reply on HN