Live data from Hacker News

AWS Post-Mortem

aws.amazon.com

21–30 of 69 posts

Re: AWS Post-Mortem

#21
post #14

Earlier quoted context omitted.

The cheapest zone, yes, so presumably the most popular. And who doesn't use EBS? You sound like you think I'm being unfair. Why? Plenty of sites were badly affected by that outage; it precipitated a lot of self-examination at some companies I know of. I don't think I am overstating the issues here.

The four zones in US-East have exactly the same prices. And I don't use EBS.

You mean availability zone. Well I guess me, Reddit, Foursquare, and plenty of other sites just got lucky in the bad availability zone.

Ah yes, here's the classic AWS apologist pattern in full effect. You don't use EBS! Of course you don't, you would have to be some kind of friggin' idiot to use EBS. So what do you use for, say, MongoDB data files, that is different from morons like me who stupidly assumed they could/should use EBS?

Re: AWS Post-Mortem

#22
post #10

Earlier quoted context omitted.

Hm. Well, I don't like them. It's subjective, you might disagree. But off the top of my head: 1. Contracts. They want 1 years minimum contracts for any dedicated servers. For truly gargantuan orders I could understand this but for one puny server? Never. 2. Their definition of "cloud" is different from mine. To use their "cloud" services your servers need to be public facing, ie on public IPs. Want them on your own V…

It sounds like you're not a fan of Rackspace (and that's fine) but you can't honestly believe that an API sends an email to a tech. That's just completely false statement that no one in there right mind should believe.

I came up with that theory because I couldn't think of anything else that explained my 2 hour waits to get a new instance. Not the other way round.

Re: AWS Post-Mortem

#23
post #17

Earlier quoted context omitted.

Is there any reason in particular why you wouldn't recommend Rackspace?

For what its worth, I found Rackspace's support to not be worth the extremely high cost. Softlayer is just as good.

Rackspace support has plummeted horribly from a few years ago. They still charge a huge premium, and act like prima donnas, but they are just not worth it.

I didn't mention softlayer because I did not want to look like a stooge but yeah, that is who I use these days. Decent prices, reasonable support, you would have to give me a reason not to go with them for a new deploy.

Re: AWS Post-Mortem

#24
post #11

Earlier quoted context omitted.

Huh? North Virginia was down for something like 30 hours in April. I had some instances in there. Hardly "incredible luck". And it was not just me: http://www.google.com.au/search?q=AWS+Virginia+outage

One zone in US-East had a significant outage in April; and only instances using EBS were affected.

I distinctly remember the Amazon status page (which unfortunately seems to have very little history) claiming that two availability zones were affected during that outage: one was fixed within two hours, and the other remained offline for over a day (and if you were as unlucky as me, was not fully recovered for multiple days, not just 30 hours).

Regardless, you have missed the point of this thread: I have a bunch of servers on standard non-cloud providers (like 1&1), and I have seriously had /one/ outage in the last EIGHT YEARS. In that case, a hard drive (just mine in one server, not half of the known Internet's) failed, was replaced, and my server was back up in a few hours.

Re: AWS Post-Mortem

#25
post #10

Earlier quoted context omitted.

Is there any reason in particular why you wouldn't recommend Rackspace?

Hm. Well, I don't like them. It's subjective, you might disagree. But off the top of my head: 1. Contracts. They want 1 years minimum contracts for any dedicated servers. For truly gargantuan orders I could understand this but for one puny server? Never. 2. Their definition of "cloud" is different from mine. To use their "cloud" services your servers need to be public facing, ie on public IPs. Want them on your own V…

(disclaimer, I work at Rackspace via the Cloudkick acquisition)

re: 2) There is a product called RackConnect, which can bridge the public cloud which has API based provisioning, and physical servers: http://www.rackspace.com/hosting_solutions/hybrid_hosting/

re 3) I'm not sure specifically what happened in your case, but generally servers are provisioned and online in the public cloud within a few minutes -- If I had to guess, its possible the huddle you were in had some kind of capacity issue or other fault that prevented immediate provisioning. If you ever boot a server and its not online in <5 minutes, I'd go straight to support chat, and they generally can tell you what is going on.

Re: AWS Post-Mortem

#26
post #4

Earlier quoted context omitted.

Sure. Presumably Amazon has a test lab that replicates multiple zones :) Perhaps your point is that Amazon should make this test lab public so people can contribute to the QA effort? IIRC many of these datacenter failures start with a utility company power outage followed by a failure of the secondary power systems (I'm thinking of some past failures at softlayer and other providers). I wonder if it is prohibitively…

I wonder if it is prohibitively expensive to do a real life system test on a big data center It's probably prohibitively dangerous. Backup power systems don't have many-nines of reliability; generators which are reliable enough for the once-a-decade event when a car crash knocks out your utility power aren't anywhere near the reliability needed to run your datacentre for an hour every month as a test.

Actually, if you don't run test your generator regularly it's very unlikely to work when you do need it.

Here's a doc from cummins, a generator mfgr: http://www.cumminspower.com/www/literature/technicalpapers/P...

It claims that the generator should be run for 30 minutes every month, loaded to at least one third of the rated capacity. So testing every month is exactly what you want to do.

Re: AWS Post-Mortem

#27
post #23
post #17

Earlier quoted context omitted.

For what its worth, I found Rackspace's support to not be worth the extremely high cost. Softlayer is just as good.

Rackspace support has plummeted horribly from a few years ago. They still charge a huge premium, and act like prima donnas, but they are just not worth it. I didn't mention softlayer because I did not want to look like a stooge but yeah, that is who I use these days. Decent prices, reasonable support, you would have to give me a reason not to go with them for a new deploy.

I also love softlayer; they have a variety of datacenters now (including san jose), great pricing (except for RAM, but you can negotiate that), and their service, while a lot smaller team than rackspace, is still very good.

Their cloud product isn't as good as Rackspace and nowhere near as good as EC2, but maybe that will improve. I still like rackspace too; I wouldn't switch from one to the other, but I would definitely evaluate both whenever making a decision about hosting for a new project.

Re: AWS Post-Mortem

#28
post #26

Earlier quoted context omitted.

I wonder if it is prohibitively expensive to do a real life system test on a big data center It's probably prohibitively dangerous. Backup power systems don't have many-nines of reliability; generators which are reliable enough for the once-a-decade event when a car crash knocks out your utility power aren't anywhere near the reliability needed to run your datacentre for an hour every month as a test.

Actually, if you don't run test your generator regularly it's very unlikely to work when you do need it. Here's a doc from cummins, a generator mfgr: http://www.cumminspower.com/www/literature/technicalpapers/P... It claims that the generator should be run for 30 minutes every month, loaded to at least one third of the rated capacity. So testing every month is exactly what you want to do.

Right, but the thing you don't test is the transfer switch/sync gear.

Powering up the generator and dumping the output as heat weekly is pretty standard practice.

Re: AWS Post-Mortem

#29
The thing I wonder about is wtf they didn't manually switch to generator when their automatic controls failed. They had presumably ~5 minutes of UPS; it took them 40 minutes to do this. This probably isn't directly Amazon's fault, but whatever contract datacenter they are using in Europe (probably a PTT, or possibly an international carrier; really curious what facility)

I'm wary of using >1 generators to back up loads, thus requiring sync on generators for backup anyway -- much more comfortable with splitting the load up by room and having one generator per, with some kind of switch to allow for pulling generators out for maintenance. This pretty much limits you to 2-3MW per room (the largest economical diesel gensets), but that's not horrible.

Really high reliability sites actually run onsite generation as PRIMARY (since it's less reliable to start), and then utility as backup. With the right onsite generation equipment, it can be cheaper/more efficient than the grid, too (by using combined cycle; use heat output to run cooling directly).

Still, the 365 Main power outages take the cake; they used rotational UPSes (generators with huge flywheels) which had software bugs such that if input power got turned off and on several times (a common utility failure mode), the unit shut itself off entirely. Doh.

Re: AWS Post-Mortem

#30
post #11

Earlier quoted context omitted.

Huh? North Virginia was down for something like 30 hours in April. I had some instances in there. Hardly "incredible luck". And it was not just me: http://www.google.com.au/search?q=AWS+Virginia+outage

One zone in US-East had a significant outage in April; and only instances using EBS were affected.

Actually some of the issues extended to the entire East region. In particular, for quite some time I couldn't create a new instance in any East zone. At all.
Post reply on HN