Live data from Hacker News

SRE Fundamentals: SLIs, SLAs and SLOs

cloudplatform.googleblog.com

61–70 of 86 posts

Re: SRE Fundamentals: SLIs, SLAs and SLOs

#61
post #57
post #54

Google: "An SLA normally involves a promise to someone using your service that its availability should meet a certain level over a certain period, and if it fails to do so then some kind of penalty will be paid. This might be a partial refund of the service subscription fee paid by customers for that period, or additional subscription time added for free." "Partial refund". That's a very low standard for a service le…

> "Partial refund". That's a very low standard for a service level agreement, but typical of Google. It seems to be the standard. The most generous SLA I've seen is 5% off the monthly bill for each 30 minutes of downtime (up to 100%). If I'm down for 10 hours, waiving one month of bills doesn't come close to the damage done. An SLA seems to be more of a promise than an agreement, because if the service goes down you'…

I've worked for a cloud provider who paid 45x for downtime. If you were down for an hour, you got 45 hours of credit on your bill.

My current ISP credits 5x - I was impacted in an outage expected to last all day, and they credited me 5 days on my next bill.

Re: SRE Fundamentals: SLIs, SLAs and SLOs

#62
post #60
post #35

Earlier quoted context omitted.

I have no idea why you’re being downvoted. It’s the same thing as Borg/Kubernetes, MapReduce/Hadoop: some things just don’t apply or aren’t as effective unless you’re operating at a huge scale and with Google’s culture.

> unless you’re operating at a huge scale and with Google’s culture. I'm not sure one has to go to the extreme of huge scale, anywhere near where Google is now, (not that that's what you said), nor all the aspects of their culture, but I agree that key fundamental aspects are often missed. My favorite example is to point out that Google does not run Hadoop on expensive, virtualized AWS instances (or even brand-name s…

[deleted]

Re: SRE Fundamentals: SLIs, SLAs and SLOs

#63
post #54

Google: "An SLA normally involves a promise to someone using your service that its availability should meet a certain level over a certain period, and if it fails to do so then some kind of penalty will be paid. This might be a partial refund of the service subscription fee paid by customers for that period, or additional subscription time added for free." "Partial refund". That's a very low standard for a service le…

This doesn't make any sense to me. How is Google supposed to be in a position to price the business risk of individual customers into a standard SLA that they offer to all their customers? That would require Google to charge different amounts of money per customer (commensurate with the business risk placed on Google's services for that customer), running actuarial numbers to ensure that Google would have the means to pay out when the SLA is violated. Doing so would place undue burden on customers, who would need to prove business risk before buying the service, and many customers are unaware of the real business risk of downtime (having not run the numbers) anyway.

With that said... Maybe it's a good product idea, to sell varying levels of SLA violation insurance alongside the service covered by the SLA. The default, free level of insurance covers the cost of the service itself, as it does today, but perhaps a customer could buy premium insurance from Google that the SLA will not be violated, increasing the payouts to offset business risk. After all, who better to put a price on the risk than Google themselves? So probably, Google can offer a better price on offsetting the risk, than a third party insurer which doesn't have access to Google's internal data.

Re: SRE Fundamentals: SLIs, SLAs and SLOs

#64
post #54

Google: "An SLA normally involves a promise to someone using your service that its availability should meet a certain level over a certain period, and if it fails to do so then some kind of penalty will be paid. This might be a partial refund of the service subscription fee paid by customers for that period, or additional subscription time added for free." "Partial refund". That's a very low standard for a service le…

I'm not aware of any SLA from any cloud provider or ISPs that offer anything other than a partial refund and/or credit. This is most certainly not specific to Google.

Re: SRE Fundamentals: SLIs, SLAs and SLOs

#65
post #11

So, how do you choose that service level objective? How do you know which solutions to implement to not make things "overly reliable"? Isn't that more important question? As doing this without some sort of methodology will almost always result in useless solutions and overpaying to cloud and other hosting providers. Like implementing rather expensive failover within the datacenter, while ignoring how unreliable datac…

Disclaimer: I am an SRE at Google, opinions are my own.

There's an excellent talk by Google VP of SRE Ben Treynor: https://www.youtube.com/watch?v=iF9NoqYBb4U. tl;dw: try to measure actual user experience, and make sure that even the long tile of customer still gets a good product experience. What "good product experience" means depends, on your product.

The rest of the error budget is for you to spend on releasing new features, changing the underlying architecture, etc.

Re: SRE Fundamentals: SLIs, SLAs and SLOs

#66
post #57

Earlier quoted context omitted.

> "Partial refund". That's a very low standard for a service level agreement, but typical of Google. It seems to be the standard. The most generous SLA I've seen is 5% off the monthly bill for each 30 minutes of downtime (up to 100%). If I'm down for 10 hours, waiving one month of bills doesn't come close to the damage done. An SLA seems to be more of a promise than an agreement, because if the service goes down you'…

I've worked for a cloud provider who paid 45x for downtime. If you were down for an hour, you got 45 hours of credit on your bill. My current ISP credits 5x - I was impacted in an outage expected to last all day, and they credited me 5 days on my next bill.

450 hours is less than a month, so that sla is actually worse than the one the above poster described, at least if you're down for less than a day.

Re: SRE Fundamentals: SLIs, SLAs and SLOs

#67
post #7

When reading these articles, never forget that your company is NOT Google! If your company doesn't have a management/infrastructure/communication/skill structure that Google has, then it will be very difficult to implement these fundamentals. In many cases, an SRE is a job to save costs. If your company doesn't get its shit together and doesn't give your SREs the support it needs, then they'll hate their jobs and the…

> In many cases, an SRE is a job to save costs. This is 100% the case. I would actually argue that it is the only job of the SRE organization - hit the budgets by balancing costs of availability vs. costs of unavailability. If the org has massive budgets and general budget flexibility then it is easy. Otherwise, SREs are magicians to pull the rabbits out of a hat inventing the most awe inspiring methods/tools/hacks/w…

Yep. It's pretty disheartening, frankly.

Re: SRE Fundamentals: SLIs, SLAs and SLOs

#68
post #54

Google: "An SLA normally involves a promise to someone using your service that its availability should meet a certain level over a certain period, and if it fails to do so then some kind of penalty will be paid. This might be a partial refund of the service subscription fee paid by customers for that period, or additional subscription time added for free." "Partial refund". That's a very low standard for a service le…

I'm not aware of any SLA from any cloud provider or ISPs that offer anything other than a partial refund and/or credit. This is most certainly not specific to Google.

Check out the contract for a lottery or gambling system provider. They usually provide that the service provider is responsible for all losses for downtime or other errors on the provider's part, including fraud and theft. GTech pays about 0.5% of their revenue in penalties.

Re: SRE Fundamentals: SLIs, SLAs and SLOs

#69
post #35
post #7

When reading these articles, never forget that your company is NOT Google! If your company doesn't have a management/infrastructure/communication/skill structure that Google has, then it will be very difficult to implement these fundamentals. In many cases, an SRE is a job to save costs. If your company doesn't get its shit together and doesn't give your SREs the support it needs, then they'll hate their jobs and the…

I have no idea why you’re being downvoted. It’s the same thing as Borg/Kubernetes, MapReduce/Hadoop: some things just don’t apply or aren’t as effective unless you’re operating at a huge scale and with Google’s culture.

No idea, either. This whole thread is being bombarded with downvotes.

Downvoters: Whatsup?

Re: SRE Fundamentals: SLIs, SLAs and SLOs

#70
post #63
post #54

Google: "An SLA normally involves a promise to someone using your service that its availability should meet a certain level over a certain period, and if it fails to do so then some kind of penalty will be paid. This might be a partial refund of the service subscription fee paid by customers for that period, or additional subscription time added for free." "Partial refund". That's a very low standard for a service le…

This doesn't make any sense to me. How is Google supposed to be in a position to price the business risk of individual customers into a standard SLA that they offer to all their customers? That would require Google to charge different amounts of money per customer (commensurate with the business risk placed on Google's services for that customer), running actuarial numbers to ensure that Google would have the means t…

The evil part of outages is that, no matter how much resource you dumped into developments toward a more reliable system, it still happens. This is true for every company including Google. So when one company is choosing between cloud providers, they compare these SLAs with themselves. Usually it's pretty hard for a random shop to reach good SLAs. So I don't see "business risk" here. Risks present all the time, CTO should try hard to minimize them but no way to remove them.

Selling insurance for SLAs seems to be an interesting idea, but this kind of insurance might be really similar to earthquake insurance, since violation of SLAs tend to be not common (otherwise why committing) but it might be a huge cascade failure once happens. Would you like to buy one? Earthquake insurance quirks all apply.

On the other side, Google has zero incentives to violate SLAs. A. You really cannot control how large the violation would be. B. Damage to branding >>>>>> money payout.

Post reply on HN