Live data from Hacker News

Client asks for 100% uptime

serverfault.com

91–100 of 100 posts

Re: Client asks for 100% uptime

#91
post #81

Earlier quoted context omitted.

Maybe it would be good to write the custom an email explaining stuff like 99%=Well run server 99.9%=Multiple backups, will cover most hardware failures 99.99%=Top grade commercial 99.999%=What companies like Google or Yahoo can achieve 99.9999%=Hopefully the US strategic defense systems are this reliable Also, you're assumption about the servers are entirely independant. That's a reasonable assumption in terms of fir…

99.999%=What companies like Google or Yahoo can achieve Even worse. Five nines is five minutes of downtime per year. The core Google search experience blew five nines for half a decade with just one outage -- the one where they marked the entire Internet as a malware site, which took somewhere like 40 minutes to address. This kind of thing makes me dismiss talk of nines as fetishism or sales-speak. You can say your s…

Or they've failed to understand exactly what the SLA says.

It may exclude all manner of things so that it says >99.5% but really doesn't mean what you as an engineer might think. These are legal agreements...

Re: Client asks for 100% uptime

#92
post #55

Earlier quoted context omitted.

I appreciate the snark in your post :), but it also brings up a serious question: How does contract law handle significant figures?

I haven't read any cases on this, but I imagine most courts would truncate/round down the achieved performance number. If the contract says 100%, that wouldn't allow for any downtime whatsoever. It might be different for lower percentages - getting 89.8% performance where 90% is called for could be a de minimis breach and not actually count as breaking the contract. Definitely curious as to whether anyone has more to…

No.

They are interpreted the same way other commercial agreements are. They're also generally very lengthy and specify what they mean (i.e. what counts, what doesn't). They also set out what the consequences of breach are (maybe you want $, maybe you want something else, etc.).

There's no magic to writing "SLA" or something else. What you put in the contract is what you'll be held to...

Re: Client asks for 100% uptime

#93
post #11

Come on, this is possible. First we're going to have to get the governments of the world together to agree to remove all nuclear weapons. Second would be getting that asteroid tracking and deflection system working. Quantum physics does unfortunately predict that the earth might flick out of existence with some small probability, but by distributing the website across the universe we can reduce this probability arbit…

Unfortunately the first law of thermodynamics says that, eventually, the sun is going to die.

Tangentially related:

http://www.qwantz.com/index.php?comic=2033

With an infinite business lifespan, the chance your service goes down due to ANYTHING rises to one.

Re: Client asks for 100% uptime

#94
post #90

Earlier quoted context omitted.

99.999% is the standard for landline telephones, which I think is a pretty decent analogy here. Unless the customer has some crummy VoIP solution and doesn't trust their phones anymore :P.

It is? According to whom? One good rainstorm can knock phone lines out for long enough to mess up five nines for a LONG time.

I believe it's what they design for at the central switch nodes. The buildings would have entire rooms if not floors dedicated to nothing but 48 volt wet-cell batteries.

Of course the "last mile" infrastructure is not so reliable. That said, in over 20 years I've had POTS service, I can't ever remember not having dial tone when I lifted the handset.

Re: Client asks for 100% uptime

#95
Let's say I wanted 100% reliable music listening. To do this, I buy a million of the original 30GB Zune media players, create a perfect failover system, so that if the sound from one of those stops for whatever reason (hardware, software, cosmic rays, etc), it'll switch to another one. I even move these Zunes all across the world, with AC provided, and multiple network links linking all of them, with satellite link backups between them.

Then December 31, 2008 rolls around, and a tiny firmware bug knocks out all of them simultaneously for 24 hours. Oops.

Not all failures are independent events.

Re: Client asks for 100% uptime

#96
post #94
post #90

Earlier quoted context omitted.

It is? According to whom? One good rainstorm can knock phone lines out for long enough to mess up five nines for a LONG time.

I believe it's what they design for at the central switch nodes. The buildings would have entire rooms if not floors dedicated to nothing but 48 volt wet-cell batteries. Of course the "last mile" infrastructure is not so reliable. That said, in over 20 years I've had POTS service, I can't ever remember not having dial tone when I lifted the handset.

Not noticing downtime is not the same as not having it.

I'll sell you 100% uptime at a much better price if you promise to only check whether you're up a few times a day...

Re: Client asks for 100% uptime

#97
post #42

I don't understand what the issue is. The client wants you to plan for disaster, and they aren't math oriented, so asking for 100% probability sounds reasonable. The engineer, as engineers are prone to do, remembered his first day of prob&stat 101, without considering that the client might not. When they say this, they aren't thinking about nuclear winter, they are thinking about Fred dumping his coffee on the office…

Most clients don't understand the R2D2 talk. They understand money, features, bugs, and downtime. So I've always explained it like so: Uptime beyond 95% costs lots and lots of money. Magnitudes of order of more money. It requires redundant equipment, engineering all of the automatic failovers at every layer, lots of monitoring, and 24/7 technical staff to watch everything like a hawk. Not ... in ... your ... budget.…

Now, you are right that you get diminishing returns as you add more nines, but 95% is still in the area where there are a lot of cheap things you can to to increase uptime; RAID, a ups, and even a low-end but business-grade connection should get you around two nines.

I have a SLA of 99.5% (over a month) on a low-end setup[1] and it's fairly rare that I don't meet it, even including planned downtime and network outages due to DoS or mistakes of my upstream.

[1]I use RAID and mostly supermicro server grade hardware with ecc ram, but there is no failover across servers; I'm in a data center, though it's a low-end data center with a low-end bandwidth provider.

Re: Client asks for 100% uptime

#98
post #81

Earlier quoted context omitted.

Maybe it would be good to write the custom an email explaining stuff like 99%=Well run server 99.9%=Multiple backups, will cover most hardware failures 99.99%=Top grade commercial 99.999%=What companies like Google or Yahoo can achieve 99.9999%=Hopefully the US strategic defense systems are this reliable Also, you're assumption about the servers are entirely independant. That's a reasonable assumption in terms of fir…

99.999%=What companies like Google or Yahoo can achieve Even worse. Five nines is five minutes of downtime per year. The core Google search experience blew five nines for half a decade with just one outage -- the one where they marked the entire Internet as a malware site, which took somewhere like 40 minutes to address. This kind of thing makes me dismiss talk of nines as fetishism or sales-speak. You can say your s…

If we're talking about the agreement, SLAs that are better than 99.9% are quite common, and available even on low end products. The problem is that the SLA payout is usually "we refund you for the time you were down, if you ask for the refund" - heck, with that payout, I'd be happy to give you a 100% SLA on any product I sell. (of course, I'm not going to advertise as such; the sort of people who buy from me would find that disingenuous.)

That said, I think you are about right with 99.5% being about the best you can expect while spending a reasonable amount of money. (especially for a static site, i think a few more tenths of a point is possible for less money than you think, but the cost curve goes parabolic sometime after 99.5%)

Re: Client asks for 100% uptime

#99
post #56
post #42

Earlier quoted context omitted.

Most clients don't understand the R2D2 talk. They understand money, features, bugs, and downtime. So I've always explained it like so: Uptime beyond 95% costs lots and lots of money. Magnitudes of order of more money. It requires redundant equipment, engineering all of the automatic failovers at every layer, lots of monitoring, and 24/7 technical staff to watch everything like a hawk. Not ... in ... your ... budget.…

First of all, uptime of 95% means 18 and a quarter entire days of downtime per year. That's horrendous. I wouldn't host my dog's website on a server with that kind of SLA - and I don't even have a dog. Secondly, although Twitter got away with large helpings of downtime, that doesn't mean that every business type can. Twitter is not (or at least was not, for most of its existence) business-critical to anyone. If Twitt…

> First of all, uptime of 95% means 18 and a quarter entire days of downtime per year. That's horrendous.

It may be acceptable to some clients depending on what other provisions are part of the SLA (though probably not with as little as 95%).

I've seen a 98% SLA which was applied both annually (~7-ana-third days) and daily (2% of a day being about half an hour) with significant remuneration if the daily SLA was not kept as well as the annual one. If I remember rightly, maintenance windows counted against the SLA except in certain (specified in the contract) circumstances.

Of course for many applications this would still be completely unacceptable, but for others it might be fine depending on the costs and the comeback if the SLA is broken.

Re: Client asks for 100% uptime

#100
post #23

Earlier quoted context omitted.

This is the kind of service I'd love to attack as a side project, it's fascinating. Though I'm sure someone out there reading this, has something like this service but affordable for startups?

There is a lot of room in this market for competition from open source projects. Really, the concept is so simple that it could be done with a shell script for simple failover: (pseudo code) if (curl http://ip of main site) fails then copy alternate zone file to bind dir service named reload fi You get the idea... F5 Networks is really just a fancy DNS server running on a BSD based OS on an x86 appliance. Zeus, which…

>I'd love to see some open source competition for this space, or even low price competition.

Yes, that's exactly what I meant. Running our own failover system is not just expensive, but time-consuming - just another thing to do when you are trying to scale and time is already short within a small team.

Post reply on HN