Live data from Hacker News

Amazon Web Services are down

status.aws.amazon.com

141–150 of 346 posts

Re: Amazon Web Services are down

#141
post #101

Assuming the problem is indeed with EBS, I would say this should be a warning sign to anyone considering going with a PaaS provider, which Amazon is quickly becoming, instead of an IaaS provider like Slicehost or Linode. The increased complexity of their offering makes it more likely that things will break, leaving you locked in. I did a 15 minute talk on the subject, which you can check out here: http://iforum.com.u…

Every time someone makes the claim that downtime should be a warning sign about going with a PaaS provider (or, indeed, an IaaS provider, or in some cases, people even make this claim about going with someone else's Data Center) - I always respond: "And why do you believe that you would do any better?" Every environment I've been involved in as an operations professional for the last 15 years has experienced downtime…

Yep.

The thing to remember is that even if you could do better -- and it is certainly possible -- your customers might not pay you for it. It will cost you, a lot, to build a significantly better system than AWS. Reliability is about redundancy, and redundancy is synonymous with paying for things that sit idle for years at a time, and paying for the hardware and personnel to run disaster drills over and over, and customers hate that. Reliability is also about avoiding excessive complexity and limiting the rate at which you change things, and customers hate that too. If they wanted highly reliable, slow-changing established tech they'd be using land-line phones, not Quora or Reddit.

Re: Amazon Web Services are down

#142
post #107
post #59

Earlier quoted context omitted.

Go back to the building metaphor. Building datacenters is not our core business, so we outsource it. Much like building a building is not our core business, nor running a telephone system, so we outsource those things too. All business critical, but still outsourced.

No, building data centers is not your core business, but your core business relies on them so intimately that it makes sense (at least to some people) to have a little more direct control of your datacenter resources. Or, the have a warm-standby of some sort. The best analogy I can make is it's like a telemarketing company. Their core business is not building/owning/maintaining PBXs, it's calling you and trying to se…

At the end of the day, you need to make money. The question becomes: can you afford to be in the datacenter business?

Executing well in operating a facility with significant scale means that you need to have folks working for you who have a clue about operating a datacenter. That's not cheap.

IMO, the real question is whether or not you should be running your core business on a genericized "public cloud" like Amazon or if you should be using a traditional server hosting outfit, a managed service provider like Verizon/AT&T or a colo. There's a value calculation that you have to compute to figure out what's right for you.

I worked for a company that ran call centers and had a significant commerce site (whose growth exploded beyond expectations) back in the early 2000's. Our offices were a converted school and the datacenter was a classroom with a roof A/C unit. We knew it sucked and suffered through downtime and failures, but the company wasn't big enough to do things the "right" way. So we had to do things the "wrong" way, because the alternative was to go out of business. If that were today, we would almost certainly have had 80% of our systems at Amazon/Rackspace facilities.

Re: Amazon Web Services are down

#143
post #126

Earlier quoted context omitted.

Every time someone makes the claim that downtime should be a warning sign about going with a PaaS provider (or, indeed, an IaaS provider, or in some cases, people even make this claim about going with someone else's Data Center) - I always respond: "And why do you believe that you would do any better?" Every environment I've been involved in as an operations professional for the last 15 years has experienced downtime…

The problem with Amazon is that despite touting an open API, their infrastructure internals and practices are a trade secret, so the likes of Eucalyptus are having to play catchup. In other words, I cannot replicate their infrastructure in my own data center, even I had the money to pay them. I suspect this is the main reason that Heroku didn't move off Amazon, and not the fact that Amazon was providing them great va…

Compared to PaaS provides, Amazon is easy to migrate off of because they are simply giving you virtual hardware.

The basics of their setup are well known: Xen and Linux. EBS is some sort of block-based network storage. True, we do not know the backend storage setup, but it does not really matter. The Xen VM simply sees a disk device. NetApp, EMC, Dell, HP and many others all have products that offer similar functionality, including snapshots.

The only part that is potentially hard to migrate off is the security groups, and then only if you are using named groups. If you are just using IP based rules, most any firewall would work.

Re: Amazon Web Services are down

#144
post #132

Earlier quoted context omitted.

Reddit's been down for several hours today, I'm sure they are already way lower 99.95%.

0.05% of one year is 4 hours, 22 minutes and 48 seconds.

Which means their entire quota for this year is all gone.

Re: Amazon Web Services are down

#145
post #2

Current status: bad things are happening in the North Virginia datacenter. EC2, EBS and RDS are all down on US-east-1. Edit: Heroku, Foursquare, Quora and Reddit are all experiencing subsequent issues.

Foursquare is up now. Quora not showing the 503 anymore, Reddit still down.

Re: Amazon Web Services are down

#146
post #132

Earlier quoted context omitted.

Reddit's been down for several hours today, I'm sure they are already way lower 99.95%.

If the Reddit web server admins took availability seriously they would have chosen to deploy across more than one region. Do you disagree? Why do you disagree? I'm being honest, no snark involved in my questions.

The whole point of AWS is to forget about maintaining hardware infrastructure.

Amazon are the ones who should have made backups in multiple regions, and transfer the load on failure.

Re: Amazon Web Services are down

#147
post #132

Earlier quoted context omitted.

Reddit's been down for several hours today, I'm sure they are already way lower 99.95%.

If the Reddit web server admins took availability seriously they would have chosen to deploy across more than one region. Do you disagree? Why do you disagree? I'm being honest, no snark involved in my questions.

Not qualified to speak about what Reddit should or should not do about the arrangement with Amazon. I have read several posts, including one by an (ex) Reddit employee saying Amazon is not delivering what they said they would, that much is clear. I really doubt all their downtime is part of the SLA.

Re: Amazon Web Services are down

#148
post #106

Earlier quoted context omitted.

I did the math on it, with the downtime so far it's almost approaching the magic 99.95% barrier where everyone gets a 10% bill credit. Wouldn't that be something to light a fire under the collective arses of those in charge of keeping EBS stable.

> Wouldn't that be something to light a fire under the collective arses of those in charge of keeping EBS stable. Yeah, the only possible explanation is that Amazon's put a bunch of lazy guys in charge of EBS. How hard a job could it be?

I'm certain they're plenty talented and hard-working, but there's got to be something in their way. Just like the guy said as he was leaving Reddit, management seems to be pointing fingers in different directions as to why this is going on and how it can be resolved. It's sad, though, because I don't want to think of Amazon as a company with a potentially toxic corporate culture. I enjoy my free Prime membership as a student :D

Re: Amazon Web Services are down

#149
post #132

Earlier quoted context omitted.

Reddit's been down for several hours today, I'm sure they are already way lower 99.95%.

0.05% of one year is 4 hours, 22 minutes and 48 seconds.

Don't forget to take in to account ALL of Reddit's downtime; there is quite a bit of it.

Re: Amazon Web Services are down

#150

Earlier quoted context omitted.

If the Reddit web server admins took availability seriously they would have chosen to deploy across more than one region. Do you disagree? Why do you disagree? I'm being honest, no snark involved in my questions.

The whole point of AWS is to forget about maintaining hardware infrastructure. Amazon are the ones who should have made backups in multiple regions, and transfer the load on failure.

"[S]hould" is the wrong word here. Clearly, they don't maintain such backups. This is clear to anyone using their service. They pay for the service anyway, so apparently it's still worth it to them, even without auto-backups.

Would it make sense for Amazon to maintain automatic backups (and potentially charge more for them)? I don't know. It might make business sense, it might not. But their service is apparently popular enough even without it.

Post reply on HN