Live data from Hacker News

Breaking Up with On-Call

reflector.dev

31–40 of 202 posts

Re: Breaking Up with On-Call

#31
post #3

From my experience working on SaaS, and improving ops at large organizations, I've seen that "on-call culture" often exists inversely proportional to incentive alignment. When engineers bear the full financial consequences of 3AM pages, they're more likely to make systems more resilient by adding graceful failure modes. When incident response becomes an organizational checkbox divorced from financial outcomes and pla…

This assumes that the engineers in question get to choose how to allot their time, and are _allowed_ to spend time to add graceful failure modes. I cannot tell you how many stories I have heard of, and companies I have directly worked at, where this power is not granted to engineers, and they are instead directed to "stop working on technical debt, we'll make time to come back to that later". Of course, time is never…

Definitely an issue but I think there's a little room for push back. Work done outside normal working hours is automatically the highest priority, by definition. It's helpful to remind people of that.

If it's important enough to deserve a page, it's top priority work. The reverse is also true (if a page isn't top priority, disable the paging alert and stick it on a dashboard or periodic checklist)

Re: Breaking Up with On-Call

#32
post #14

A moderate on-call ritual is a necessary evil. I’ve worked at places that tried to get rid of it with all kinds of automation and playbooks, only to revert back to PagerDuty a few months later. That said, my last workplace completely burned me out with a terrible on-call policy and an absurdly short recovery period. Not to mention, upper management tried to gaslight everyone into thinking on-call was just part of nor…

> Not to mention, upper management tried to gaslight everyone into thinking on-call was just part of normal work and didn’t warrant additional compensation. Every position I've held (save my current one), that is most definitely the norm. If you have a busy night managers are cool with you coming in late the next day (or potentially not at all), but it's very unusual to be paid for on call in my experience.

Fortunately every job I've held had the situation where the manager was fine with me/the team taking make-up time the following day if you were interrupted during offhours. Sure extra comp would be nice, but if on-call has to be done, this isn't the worst way.

Re: Breaking Up with On-Call

#33

Been a few years since I worked at Google as an SRE, but I did not find "There’s no incentive in big tech to write software with no bugs" to be particularly true. Perhaps because I was an SRE for a few of the of the older, intrinsic, core products. A lot of the pages we got were for things outside our control (e.g. some fiber optic cables got broken (we found out later), so we've got to drain some cluster because not…

Were the systems designed to scale based on load and handle transient failures?

It seems like lack of automated remediation would be a bug unless it's an "accepted business risk" i.e. cheaper to throw people at to manually fix than build a software solution.

Re: Breaking Up with On-Call

#34

Earlier quoted context omitted.

great points. I think over the 8 years of my SRE experience, I've probably caused a few outages after being paged at 3-4am and/or prolonged them. I've fallen asleep during one responding time (fortunately the main outage had been fixed but we went way beyond our internal SLO for backups as I had fallen asleep before running them). That said, the author neglected to mention timezones or following the sun. A 12/12 hour…

> That said, the author neglected to mention timezones or following the sun. A 12/12 hour shift or 8/8/8 (NA/EU, NA/EU/EMEA) addresses the sleep deprivation problem pretty well I completely agree. The "get a call at 3am" scenario, to me, is shorthand for an organization which intentionally under-staffs in order to save money. If a system has genuine 24x7 support needs, be it SLA's or inferior construction, it is incu…

It always amazes me these places that have "24/7 support needs" then all of a sudden a bug comes in that's "not important" even though it has customer uptime impacts.

Re: Breaking Up with On-Call

#35
post #14

A moderate on-call ritual is a necessary evil. I’ve worked at places that tried to get rid of it with all kinds of automation and playbooks, only to revert back to PagerDuty a few months later. That said, my last workplace completely burned me out with a terrible on-call policy and an absurdly short recovery period. Not to mention, upper management tried to gaslight everyone into thinking on-call was just part of nor…

I don't think on-call is a necessary evil; I think it's the result of managers and leaders not caring about having multiple failsafes instead opting to foist that problem onto engineers via unpaid labor.

You can have a system with enough redundancy, ability to rollback and deployment scheduling where any sort of on-call incident is highly rare and low impact. But that requires spending time and money on solving those problems which is money that could be spent on developing more features faster. A buggy and broken product doesn't matter as long as you can shovel out More faster.

I've been on both types of teams and I'm at the point in my career that if I'm going to be on-call I'm expecting to be compensated appropriately with actual on-call hours worked. And part of the problem is that even on a team where you're informed of on-call duties and rotation, if the team gets cut or people leave then you're on the hook for working longer hours for less pay. It's inherently exploitative.

Re: Breaking Up with On-Call

#36
> Filing for a new feature implementation would require a thorough documentation, rightfully so, followed by a political campaign to convince the political party of principal engineers and managers to accept the new feature. These stakeholders carry incentives and principals of their own - that do not necessarily always align with the true engineering spirit of solving the problem, nor with satisfying the customer.

> The same friction applied to fixing a bug or a flawed process. I would reproduce the bug: spin up the entire environment, the appropriate binary artifact and the reproduced state of application, create the test cases and pin-point the exact problem for the stakeholders as well as present the possible solutions. I would get sent to a dozen of meetings, bouncing my ideas back and forth until receiving the dire verdict - rejection to fix this bug altogether.

This has been my experience for the entire duration of my tenure at the current place I am employed at. I've stopped doing this, because my backlog is filled with "best effort" features and when I attempt to slot some of these into a quiet sprint, management says no.

Many features I request is not from assumption, but data I gathered with analytics for the customer-facing website. One example is to change a page layout and add better search functionality for this page and its data, because I noticed a +70% no-results rate in the search analytics. I suggested a change but marketing shot it down saying it's low priority for them, while they frequently say that the website can perform better at generating leads. I might be wrong, but to me it feels like utter short-sightedness and goes against the strategic goal of the company.

I'm just an engineer, what do I know?

Re: Breaking Up with On-Call

#39
post #15
post #3

From my experience working on SaaS, and improving ops at large organizations, I've seen that "on-call culture" often exists inversely proportional to incentive alignment. When engineers bear the full financial consequences of 3AM pages, they're more likely to make systems more resilient by adding graceful failure modes. When incident response becomes an organizational checkbox divorced from financial outcomes and pla…

> When engineers bear the full financial consequences of 3AM pages, they're more likely to make systems more resilient by adding graceful failure modes. Making engineers handle 3 AM issues caused by their code is one thing, but making them bear the financial consequences is another. That’s how you create a blame-game culture where everyone is afraid to deploy at the end of the day or touch anything they don’t fully u…

"Financial consequences" probably mean "the success of the startup, so your options won't be worth less than the toilet paper", rather than "you'll pay for the downtime out of your salary".

Re: Breaking Up with On-Call

#40

Not to distract from the article, but I'm pretty sure that photo of the guard tower is from Manzanar, one of the concentration camps in California that interned Japanese Americans during WWII. Probably not in great taste to use that photo to represent the idea of "guard duty" in software. https://en.wikipedia.org/wiki/Manzanar

Who, specifically, would that picture offend or change protect?

People that have been interned in this or one of the other camps, or their descendants. It’s one generation ago, people born and raised in camps are still alive. George Takeo for example was born in one of the segregation camps.

It’s a low stakes change.

Nobody assumes harmful intentions from the author - I would not have recognized the picture either. But now that it’s been pointed out that it’s from a site where people were held illegally against their will, the reaction is a tell-tale. Knowing this, and insisting on keeping the images is now willfully associating with harmful behavior.

Apart from that, I cannot associate with the picture either - as an on-call engineer for some widely used infrastructure, I am not a guard on duty keeping people in a camp. I am an emergency responder. I fix things when they go haywire or an accident happens. A firefighter, paramedics, civil emergency responder would IMO be a much better metaphor for what I do.

Post reply on HN