Live data from Hacker News

Redefining Observability

hazelweakly.me

11–20 of 23 posts

Re: Redefining Observability

#11
I think this blog post uses a lot of words to not say much; maybe useful as an exercise in burying the lede?

I read the whole thing and it matches what I consider to be important in observability ("Help me make a decision"), but the point is so spread out, and so diluted that I'm not even sure that it is there at all.

So, firstly, here's two easy and simple rules for technologists trying to satisfy the audience for the observability outputs:

1. If what you get out of the observability exercise/process/tools doesn't help make a decision, that's not observability, it's trivia.

2. If you're making decisions without any data or with insufficient data, then you don't have any (or have insufficient) observability.

Depending on the audience, different people make different decisions.

Secondly, here's a cheat-sheet to help: https://www.rundata.co.za/rundata/blog/chap2/content.html

Re: Redefining Observability

#12

Earlier quoted context omitted.

I hate all of these goofy neologisms corporate culture comes up with ("impactful") that are just abuses of syntax and have no reason to exist because there's already a word that means the same thing. In this case the word visibility already means exactly what's being described here (impactful -> impressive/effective/significant).

I have to disagree with this. Yes existing terms have meaning, but they also have baggage. Forming a new term is a reasonably effective way to shed that baggage and attempt to start the cycle all over again. Seeing a term you’re not as familiar with and needing to look up the definition (whether that’s a local or a global definition) is usually a good thing IME.

This is naturally happening in all languages, but there’s a big difference between widely accepted new norm and a jargon invented by a small group of people who didn’t bother to think enough.

Re: Redefining Observability

#13
My "become the joker" take is that the SRE book is completely misunderstood and misapplied by small shops that just do not have the profit margins of GOOG to staff internal tooling & operations at the same scale.

So instead of spending scarce resources on the 80/20 monitoring of compute/memory/disk/network resources we see shops build Rube Goldberg machines to drive KPI dashboards that the CTO can stare at. Except that there's no one really spending the time to pick meaningful KPIs and levels.

Hey, I get it.. "cattle not pets".. but, if some pods are constantly falling over then maybe monitor & address that rather than waiting for some 3 steps removed KPI on error rates crosses the 97.5% level. I've been at too many shops where they got observability-pilled and constantly still fell over on the types of underlying resource starvation outages most tooling had well addressed by 2005. If you don't notice you ran out of some storage resources last night until a cascade of failures cause a user facing data availability SLA to be breached.. that's silly.

Re: Redefining Observability

#16

I think your definition of observability sits on top of the other two (Control & Cognitive definitions) and ultimately is more of a description of an approach to observability rather than a redefinition of it. Observability is hard and I agree with the approach you've laid out, especially on the involvement of leadership insight. It's crucial and it has been omitted or ignored while working on this for a decade. The…

I think I agree to some extent. I like parts of all 3 definitions presented here, and I feel the new definition loses some of the clarity by omitting the fact that we’re interested in a particular system at a time.

But I really like this blog post’s main thesis that the people are part of the system and a key part of your observability stack. You can have an objectively perfectly observable system that outputs all of its answers to your questions in Greek, and I won’t be able to make head nor tail of it and neither will pretty much any of my team. Just as nines don’t matter if the users aren’t happy, observability doesn’t matter if the team can’t understand and act on the answers.

I have spent the last year building out observability tooling and I expect to spend another year on it. I would say 70% of my problems to solve are social rather than technical.

Re: Redefining Observability

#17

My "become the joker" take is that the SRE book is completely misunderstood and misapplied by small shops that just do not have the profit margins of GOOG to staff internal tooling & operations at the same scale. So instead of spending scarce resources on the 80/20 monitoring of compute/memory/disk/network resources we see shops build Rube Goldberg machines to drive KPI dashboards that the CTO can stare at. Except th…

That's not "become the joker;" that's facts. The SRE book was, literally, Googlers documenting how they did SRE at Google. The book (i.e. the first two chapters) got cargo-culted everywhere, everything else became footnotes. See also: The massive set difference of places that have something resembling SLOs and places that implement interrupts and fair compensation for on-call rotations.

Re: Redefining Observability

#18
post #17

My "become the joker" take is that the SRE book is completely misunderstood and misapplied by small shops that just do not have the profit margins of GOOG to staff internal tooling & operations at the same scale. So instead of spending scarce resources on the 80/20 monitoring of compute/memory/disk/network resources we see shops build Rube Goldberg machines to drive KPI dashboards that the CTO can stare at. Except th…

That's not "become the joker;" that's facts. The SRE book was, literally, Googlers documenting how they did SRE at Google. The book (i.e. the first two chapters) got cargo-culted everywhere, everything else became footnotes. See also: The massive set difference of places that have something resembling SLOs and places that implement interrupts and fair compensation for on-call rotations.

Yes, exactly.

I've seen small companies that are smaller than Google's SRE staff alone try to implement the concept. As it turns out if your entire tech org is 50~200 people, hiring 2 people to "do SRE" because the CTO read a book.. is not gonna do the trick.

CTO gonna get pretty dashboard though!

Re: Redefining Observability

#19

Earlier quoted context omitted.

I hate all of these goofy neologisms corporate culture comes up with ("impactful") that are just abuses of syntax and have no reason to exist because there's already a word that means the same thing. In this case the word visibility already means exactly what's being described here (impactful -> impressive/effective/significant).

I have to disagree with this. Yes existing terms have meaning, but they also have baggage. Forming a new term is a reasonably effective way to shed that baggage and attempt to start the cycle all over again. Seeing a term you’re not as familiar with and needing to look up the definition (whether that’s a local or a global definition) is usually a good thing IME.

> needing to look up the definition

Who has ever needed to look up the definition of something like "impactful" or "observability"...?

Re: Redefining Observability

#20

Earlier quoted context omitted.

I have to disagree with this. Yes existing terms have meaning, but they also have baggage. Forming a new term is a reasonably effective way to shed that baggage and attempt to start the cycle all over again. Seeing a term you’re not as familiar with and needing to look up the definition (whether that’s a local or a global definition) is usually a good thing IME.

> needing to look up the definition Who has ever needed to look up the definition of something like "impactful" or "observability"...?

Folks who don't have English as their native language...?
Post reply on HN