Live data from Hacker News

AWS Post-Mortem

aws.amazon.com

31–40 of 69 posts

Re: AWS Post-Mortem

#31
post #21

Earlier quoted context omitted.

The four zones in US-East have exactly the same prices. And I don't use EBS.

You mean availability zone. Well I guess me, Reddit, Foursquare, and plenty of other sites just got lucky in the bad availability zone. Ah yes, here's the classic AWS apologist pattern in full effect. You don't use EBS! Of course you don't, you would have to be some kind of friggin' idiot to use EBS. So what do you use for, say, MongoDB data files, that is different from morons like me who stupidly assumed they could…

Only fools would want a persistent filesystem for their OS!

Re: AWS Post-Mortem

#32
post #21

Earlier quoted context omitted.

The four zones in US-East have exactly the same prices. And I don't use EBS.

You mean availability zone. Well I guess me, Reddit, Foursquare, and plenty of other sites just got lucky in the bad availability zone. Ah yes, here's the classic AWS apologist pattern in full effect. You don't use EBS! Of course you don't, you would have to be some kind of friggin' idiot to use EBS. So what do you use for, say, MongoDB data files, that is different from morons like me who stupidly assumed they could…

We use EBS, and had machines in the availability zone that went down that were affected. Those machines were out for longer than a day, but we were back up within an hour because we had redundancies built in across other availability zones.

If you're doing anything that matters, you can't rely on a single zone/machine/whatever, no matter who your hosting provider is.

Re: AWS Post-Mortem

#33
post #6

As someone who has put a considerable amount of resources moving things into cloud computing - I wanted to believe. But I have changed my mind. Cloud computing scales the efficiencies, yes. It also scales the problems. And because of this, AWS is by several orders of magnitude the worst of my current hosts. I have dedicated servers. No downtime in past year. I have a couple of cloud servers with rackspace. No downtim…

I'm the same. I fully embraced S3 and EC2 when they came out (even played with SQS) and enthusiastically told everyone I could that this was the future, it's the new electricity, etc.

While I still think that eventually it will end up as a utility I'm opting out of the cloud for anything production for the time being. I'll keep an eye on it of course. Does anyone know if Heroku has spread their services across HA zones?

I'll be keeping my Linodes though. They've been great.

Re: AWS Post-Mortem

#34
post #28
post #26

Earlier quoted context omitted.

Actually, if you don't run test your generator regularly it's very unlikely to work when you do need it. Here's a doc from cummins, a generator mfgr: http://www.cumminspower.com/www/literature/technicalpapers/P... It claims that the generator should be run for 30 minutes every month, loaded to at least one third of the rated capacity. So testing every month is exactly what you want to do.

Right, but the thing you don't test is the transfer switch/sync gear. Powering up the generator and dumping the output as heat weekly is pretty standard practice.

Also don't forget to check the fuel tanks. With the rise in fuel prices the past couple of years, theft of diesel from backup generators has become more common.

Re: AWS Post-Mortem

#35
post #4

Earlier quoted context omitted.

Sure. Presumably Amazon has a test lab that replicates multiple zones :) Perhaps your point is that Amazon should make this test lab public so people can contribute to the QA effort? IIRC many of these datacenter failures start with a utility company power outage followed by a failure of the secondary power systems (I'm thinking of some past failures at softlayer and other providers). I wonder if it is prohibitively…

I wonder if it is prohibitively expensive to do a real life system test on a big data center It's probably prohibitively dangerous. Backup power systems don't have many-nines of reliability; generators which are reliable enough for the once-a-decade event when a car crash knocks out your utility power aren't anywhere near the reliability needed to run your datacentre for an hour every month as a test.

In general active engines (and fuel) don't store well. They are full of lots of seals and fluids that need to be exercised periodically to function correctly. As someone else posted generator manufacturers recommend running them once a month to keep them in good working order. The same is true of a car, leave it parked in one place for too long, and you are going to have trouble starting or driving it.

Re: AWS Post-Mortem

#36
post #6

As someone who has put a considerable amount of resources moving things into cloud computing - I wanted to believe. But I have changed my mind. Cloud computing scales the efficiencies, yes. It also scales the problems. And because of this, AWS is by several orders of magnitude the worst of my current hosts. I have dedicated servers. No downtime in past year. I have a couple of cloud servers with rackspace. No downtim…

Is there any reason in particular why you wouldn't recommend Rackspace?

"As part of our efforts to continually improve our Rackspace Cloud offerings, we will be performing maintenance on our Cloud Servers environment. The maintenance window has been rescheduled and will now occur on Friday November 12th, 2010 from 10:00 pm US-CDT (3:00 am GMT) and end Saturday November 13th, 2010 at 10:00am US-CDT (3:00 pm GMT). This maintenance is required to update our billing systems. These changes will not affect nor change any of your current billing fees."

Ok, scheduled down-time is better than unscheduled down-time, but it's still down-time.

Re: AWS Post-Mortem

#37
post #31
post #21

Earlier quoted context omitted.

You mean availability zone. Well I guess me, Reddit, Foursquare, and plenty of other sites just got lucky in the bad availability zone. Ah yes, here's the classic AWS apologist pattern in full effect. You don't use EBS! Of course you don't, you would have to be some kind of friggin' idiot to use EBS. So what do you use for, say, MongoDB data files, that is different from morons like me who stupidly assumed they could…

Only fools would want a persistent filesystem for their OS!

[deleted]

Re: AWS Post-Mortem

#38
post #21

Earlier quoted context omitted.

The four zones in US-East have exactly the same prices. And I don't use EBS.

You mean availability zone. Well I guess me, Reddit, Foursquare, and plenty of other sites just got lucky in the bad availability zone. Ah yes, here's the classic AWS apologist pattern in full effect. You don't use EBS! Of course you don't, you would have to be some kind of friggin' idiot to use EBS. So what do you use for, say, MongoDB data files, that is different from morons like me who stupidly assumed they could…

S3?

Re: AWS Post-Mortem

#39
post #31
post #21

Earlier quoted context omitted.

You mean availability zone. Well I guess me, Reddit, Foursquare, and plenty of other sites just got lucky in the bad availability zone. Ah yes, here's the classic AWS apologist pattern in full effect. You don't use EBS! Of course you don't, you would have to be some kind of friggin' idiot to use EBS. So what do you use for, say, MongoDB data files, that is different from morons like me who stupidly assumed they could…

Only fools would want a persistent filesystem for their OS!

How is S3 not persistent...

Re: AWS Post-Mortem

#40
I thought best practice backup power was to use large flywheels for re-generation, and spin up diesel engines to power the wheel in the event of a loss. That way, there is no phase synchronization issue, just a mechanical clutch. Seems like this outage could have been prevented with better gear?
Post reply on HN