I like the idea of modelling availability/reliability for this. Even if you don't have the right numbers and do it on a napkin, not in code, it still can highlight solutions with best cost/benefit ratios.
SRE Fundamentals: SLIs, SLAs and SLOs
11–20 of 86 posts
Re: SRE Fundamentals: SLIs, SLAs and SLOs
#12This is an interesting article from a company that has almost nil customer support.
Re: SRE Fundamentals: SLIs, SLAs and SLOs
#13This is an interesting article from a company that has almost nil customer support.
From the movie The Negotiator: A Marine and a sailor are taking a piss. The Marine goes to leave without washing up. The sailor says, 'In the Navy they teach us to wash our hands.' The Marine turns to him and says 'in the Marines they teach us not to piss on our hands'. BTW it's not true that Google has almost nil customer support. There's extensive support for paying customers (for ads, GCP, GSuite etc.). But it's a…
Except when they aren't and even have to get data back from backup tapes.
Re: SRE Fundamentals: SLIs, SLAs and SLOs
#14In general, it's good to be precise about how you measure and when something is a hard or soft boundary. Otherwise, firefighting gets out of control. It's hard to determine when to stop something and put out a fire if you can't prioritize issues based on the boundaries you've set for your system.
Re: SRE Fundamentals: SLIs, SLAs and SLOs
#15Earlier quoted context omitted.
I have to disagree. The typical and intuitive ways of reasoning about outages and outage risk - screaming at the engineers until they fix it, desperately passing the buck, finding someone to fire in the aftermath - are not a good fit for any context. Every company can benefit from a more principled mental model of system reliability.
If your company's management doesn't even know what an SRE is, then you're stuck in the same exact place, where the SREs are the one being screamed at instead. Some companies just rename "devops" to "SRE".
Re: SRE Fundamentals: SLIs, SLAs and SLOs
#16"Excessive availability can become a problem because now it’s the expectation. Don’t make your system overly reliable if you don’t intend to commit to it to being that reliable."
Re: SRE Fundamentals: SLIs, SLAs and SLOs
#17Re: SRE Fundamentals: SLIs, SLAs and SLOs
#18This is a great article for defining terms. For some reason though, this quote made me laugh out loud: "Excessive availability can become a problem because now it’s the expectation. Don’t make your system overly reliable if you don’t intend to commit to it to being that reliable."
Re: SRE Fundamentals: SLIs, SLAs and SLOs
#19This is an interesting article from a company that has almost nil customer support.
IIRC they define their customers as other internal teams, separate from external customers, who (I'm guessing) are handled by external product teams.
Re: SRE Fundamentals: SLIs, SLAs and SLOs
#20Earlier quoted context omitted.
If your company's management doesn't even know what an SRE is, then you're stuck in the same exact place, where the SREs are the one being screamed at instead. Some companies just rename "devops" to "SRE".
I think the renaming is fine as long as it also comes with the responsibility of driving the tracking and improving of site reliability :)