Live data from Hacker News

Developer on Call

henrikwarne.com

221–230 of 246 posts

Re: Developer on Call

#221
post #48

Earlier quoted context omitted.

If you're not getting a written contract you're not working in the industry, you're the victim of a fraud.

One that discusses the work you’ll be doing, your hours/schedule, and whether or not there is on call work? I’ve never seen anything like this.

Bizarre. As a Brit I've never not had one. And I know US companies are perfectly capable of doing them for international employees.

The culture shock I'm feeling at this discovery is worse than discovering the US doesn't have electric kettles, bans crossing the road, or still uses cheques in shops.

Re: Developer on Call

#222
While I agree developers should be responsible for their work, I'm very wary of "Why Developers Should Be On Call" going the way of the whole "open-office layout". Fast spreading, but abused by many companies to optimize for the bottom line without much care for anything else.

I recently had an incredibly dystopian experience around being on-call as a developer, and while I know for a fact that's not the norm, it's enough cause for concern to share my experience with others in hopes companies that choose this are held to higher standards and processes.

I joined a company in Vancouver early this year, that I will call company X. Company X is a well known name in the U.S for real estate/property search/etc. I was hired onboard to help transition a good chunk of their dated front-end code and help champion the direction of the front-end for various product teams in the company. Turns out the front-end was a giant amalgamation of a couple things: Dust.js, jQuery, bits of really poorly written React.js, all hooked up with and plugged into Node.js rendered server-side pages. An immense amount of UI bugs and regressions would appear whenever anyone haphazardly made a change to a seemingly unrelated component/page. Multiple efforts over the years were made by various people to "take the lead" on coming up with a shared UI/component library that was to be used across the various teams and products, but the components themselves were very buggy and lacked clear, consistent design patterns or input from UX/UI designers. This caused most of the teams to resort to building their own variations of similar components, with little effort to contribute back. This would continue over a couple iterations until someone else came up with the genius idea to build a share UI/component library...you get the idea. To actually develop and make changes on the front-end was even more archaic. The various products owned by the teams occupied a portion of the site, and were all hooked up by a build harness that someone had created. Only one person really knew how the harness worked, you needed to be able to connect to a specific machine to even just load the site navigation or anything, for that matter. There was a whole week or two where this wasn't possible, and productivity slowed to a crawl. Interestingly enough, the version of the harness that various teams were running were also different and out of sync. So you'd run the harness and wait some 3 minutes to test any little change, but no other pages nor products worked, so if your feature required integration with various other products, you were in for one hell of a ride. On top of this, a lot of the front-end code was written by developers that weren't well versed in building front-ends for web applications. Needless to say, the codebase was largely an entangled mess of different ideas, state management strategies, polluting of the global namespace, front-end libraries, duplicate code, hacks, and nuances. Some 2~3 years prior to my joining, the company had a mass exodus of developers -- apparently the place is rife with political turmoil amongst various directors and departments, too.

Prior to joining, I was explicitly told there was no on-call. Some 3 or so weeks after, there was talk about "testing Pagerduty". Very quickly, every developer on the product teams were required to be hooked up to Pagerduty and be on a recurring schedule. This is what that looked like for my team: 2 developers would be on-call on any given week, for 2 straight weeks. The intern, contractor, and Principal were excluded. This meant that as 1 of the 4 other people on the team, you'd be on-call 24/7 for 2 weeks every 4 weeks. How were the escalation and notification policies setup? When any error occurred, you'd get an app notification from Pagerduty, immediately followed by a text message, and a phone call. If you did not acknowledge within 3 minutes, it would text, phone, and notify again every minute until 5 minutes. At the 5 minute mark it would call the other 2 developers. No ack in 15 minutes -> Principal + Manager, next 15 minutes -> Director. My manager had 2 teams under him, and at one point he got an escalation from his other team. Saying he was unhappy would be an understatement -- a large number of hours and meetings over the next couple weeks were put in place to come up with a plan to make sure it never happened again and to keep people accountable.

Frequency of on-call rotation and overly aggressive escalation policies aside, there were other major issues. Traditionally, the products/services were all part of one large monolithic application. At some point in the past 2 years, there was a big push towards microservices. However, there was no API versioning, no proper logging or much ability at all to track where an error originated from. Despite using microservices, deployments were a coordinated effort every Thursday, along with code freeze and multiple rungs of approval from PMs to Directors/VP. Unfortunately, the team I was on was in charge of the CRM portion of the product, which was the most commonly used feature and had many integrations with other teams. This meant that for many teams, their errors would only bubble up through our front-end, where Pagerduty would be triggered for our team. In order to make the alerts stop, there were a number of hurdles. Firstly, there was no way to snooze some of these alerts as they weren't identified as identical errors even though they were. Secondly, locating the root of the issue was often extremely difficult, between the broken build processes and fragmentation. Thirdly, as APIs weren't versioned and deployments were done once a week as a concerted effort, fixes would not land until at least the next week, at best.

There were multiple times when I was on-call that I'd be woken up multiple times at incredibly inconvenient times: 2am, 4am, 5am, any day, didn't matter. Pagerduty bombardment came frequently. One day in particular I was at my desk trying to get work done and my phone went off some 13 times in 1 hour, all first alerts, and for the same issue. The cause? One of the teams was in charge of maintaining a set of APIs around Twilio, and pushed an update that caused constant errors everytime someone made a call. Obviously, this surfaced through our team instead of theirs. There was no rollback or anything to address this immediately. After tracking down the root cause and making the team aware, they had to prioritize the issue so it could get a resolution. The fix took just over 3 weeks, during which time all our team could do was put up with the pages and dismiss them.

I'd expressed concerns around how Pagerduty would be put into place prior to all this happening, and during. Throughout, the response from management was very clear: tough luck, deal with it or get out (in more words). Multiple members on both my manager's teams (amongst other teams) expressed discontent and frustration, many talks were had, and all fell on deaf ears. To top it all off, there was zero compensation, both monetary and time off. Myself and another colleague left, yet another transferred to a different part of the company without Pagerduty, and now another mass exodus is in full swing. Even the new contractor decided to get out well before his 8 months was up.

Overall it was a horrid experience, an incredible waste of everyone's time, productivity, health, and money. I'd hate to see this type of paradigm proliferate in the industry without due diligence and care around the whole practice. All I have left to show for it is my body in a constant state of anxiety, as if I'm still on 24/7 Pagerduty.

Re: Developer on Call

#223

Earlier quoted context omitted.

So if your measure of "objective" goodness is utilitarian calculus, as it seems to be, you're leaving out the fact that employees getting the shaft tends to correlate with cheaper goods and services. I disagree that this is any more objective than my initial assessment that "this is OK for now" or my later assessment that "this sucks, I'm going back to school." But this all comes back to what I said about "objective"…

Not utilitarian calculus. It's closer to spider-man's Great power comes great responsibility. For example, if I figured out the secret to creating strong AI with respect to writing software such that I could replace the entire software engineering industry with one large computer (note: this isn't something I believe will be possible for centuries if it is ever possible), then I would feel compelled to use the billio…

It's disagreements like this that keep me coming back to HN. Thanks for a productive discussion!

Re: Developer on Call

#224
post #206

Earlier quoted context omitted.

> Does that not sound manipulative to you? You seem to be assuming there is a lot of peer pressure placed on you if you don’t want to do it. Why? I’m simply saying there are always social costs. For example you probably won’t be listened to as much when there are conversations around improving system stability. It’s like our after work Friday drinks are entirely optional - but lots of people build friendships and tru…

> For example you probably won’t be listened to as much when there are conversations around improving system stability. That sounds like not listening to people about things they might be good at and know something about, because you want to punish them for something completely unrelated. Namely, punish them for not participating in "optional" activities. All the while you don't want to openly and transparently say w…

> would listen and judge system stability suggestions based on participation in supposedly optional activity unrelated to system stability.

In my experience they are closely related.

> You also openly say that you trust people work based on Friday beer

Sure - there is an incredible depth of research on trust building via outside of work/after work activities.

> That sounds like horrible workplace

Strange considering I work at companies regularly listed in “best companies to work for” surveys.

Re: Developer on Call

#225

Earlier quoted context omitted.

They should be free to live their life. If it turns out that the company doesn't actually have enough money to hire enough people to perform all of the duties that it needs to continue to exist, then that company should cease to exist. Which is desirable over the alternative of having rich company owners externalize their failure to run their company adequately by stealing the lives of employees who don't have the ab…

> They should be free to live their life. How does an on-call rotation prevent this?

Depends entirely on the on-call rotation. I've been on one where you could go the entire week without being paged, and where most of the pages did not require an immediate response. That did not prevent me from living my life. I've been on another where there were several pages every day, at all hours of the day, each of which could take anywhere from 30 minutes to several hours. That one certainly prevented me from living my life.

Re: Developer on Call

#226

Kayak.com co-founder Paul English on this topic (2010): "The engineers and I handle customer support. When I tell people that, they look at me like I'm smoking crack. They say, "Why would you pay an engineer $150,000 to answer phones when you could pay someone in Arizona $8 an hour?" If you make the engineers answer e-mails and phone calls from the customers, the second or third time they get the same question, they'…

That's what PMs are supposed to be for.

Re: Developer on Call

#227
post #115

There's a lot of words, in the article and the comments, about being compensated for on-call time. The software field seems very anti-union, for reasons that I don't entirely understand. Protecting your time is one feature that unions offer. Want to be paid for every hour you're on call? Get the union to put it in their rules for employers. The alternative we have now is that each developer is responsible for negotia…

If Silly Valley was run like Hollywood, you'd have Version Control Engineers as a class and no one with their Version Control's Guild cards would have any say in that field. (With how Hollywood is run, you might as well have While Loop Engineers.)

Hollywood isn't the only union in the world, and from what I can tell, it's unique in its manner of operation. Is there any other trade union which has "cards" like that?

Re: Developer on Call

#228
post #140
post #115

There's a lot of words, in the article and the comments, about being compensated for on-call time. The software field seems very anti-union, for reasons that I don't entirely understand. Protecting your time is one feature that unions offer. Want to be paid for every hour you're on call? Get the union to put it in their rules for employers. The alternative we have now is that each developer is responsible for negotia…

> The software field seems very anti-union, for reasons that I don't entirely understand. A few ideas, not in any specific order: It would kill start-ups, wouldn't it? If you have to obey union rules, you can't have one person who's the DBA and the project lead and the primary programmer and and and because those are all different jobs and require different union employees. Great if you're Microsoft and can afford it…

I suggested one benefit of a union: pay for on-call hours. You've extrapolated from this at least 3 or 4 other rules that I never mentioned, and which I've never heard of, and which make no sense to me. Where is all this coming from?

Why would the union rules require separate people for two positions? Why would this be relevant to open source and free software? Why would it be relevant to hobbyists? And so on. I just don't see these problems in other fields. There's unions for actors, and stagehands, and janitors, but I still have a local amateur theatre in my neighborhood, and the lighting director sweeps the floors when that needs doing.

Whenever I say the word "union", all the complaints I hear against them from programmers sound like straw men. That makes me believe that it's the right track.

Re: Developer on Call

#229
post #154
post #115

There's a lot of words, in the article and the comments, about being compensated for on-call time. The software field seems very anti-union, for reasons that I don't entirely understand. Protecting your time is one feature that unions offer. Want to be paid for every hour you're on call? Get the union to put it in their rules for employers. The alternative we have now is that each developer is responsible for negotia…

A union is a mechanism by which a class of people with less power (typically workers) pool together to enforce rules on people with more power (typically employers). Right now, software developers have LOTS OF POWER. Companies are constantly courting devs, offering very high salaries, and including incredible perks that simply aren't available in most other industries. Yes, some engineers are exploited, but on averag…

It's a simplistic model where workers have "less power" and employers have "more power". Both sides need each other. Workers only have less power individually because they negotiate separately. That's as true for software developers as it is for anyone else.

The "perks" I see being offered for software developers are essentially free for employers to offer at scale. When you're paying someone a 6-digit salary, "free drinks and food" at work are cute but pointless. Anything of real value, like working conditions, are never up for negotiation. I see lots of people here saying they've never been paid for being on-call (nor have I). Most people I've asked say they'd prefer, and are more productive in, a private office, but no employer in my city offers it. You can't just switch employers to fix this.

"Union overhead" is very low, for what workers get out of it. Strikes are actually quite rare. Boeing machinists are a giant union here, and they haven't been on strike in over a decade. Voting and compliance are strange things to complain about. Either they're already done today, but in different contexts (e.g., one-on-one meetings, if your manager listens to you and has the power to give you what you want), or they're not done today, and it would be a great improvement for workers if it were.

These complaints sound downright bizarre in any other context. I don't hear anyone looking at a modern democratic republic with safety regulations and saying "Well, things like voting and compliance are just more trouble than they're worth".

Re: Developer on Call

#230

Earlier quoted context omitted.

An outage could be caused by infrastructure, by library code, by configuration, by increased activity, not necessarily your own software.

Yes, and there's at least two ways you could handle this: - Have the entire org take a hit - Penalize the infra team The thing about blame assignment is that no one wants to get blamed for anything (so like if you try #2 the infra team would likely find someone to blame as quick as possible), which ordinarily makes it pretty toxic but I think you can use it for good here with proper communication and goal-setting (wh…

What you want is to have fewer outages. Penalization is a non-goal.
Post reply on HN