Live data from Hacker News

Decisions that eroded trust in Azure – by a former Azure Core engineer

isolveproblems.substack.com

621–630 of 697 posts

Re: Decisions that eroded trust in Azure – by a former Azure Core engineer

#621

This read was a blast from the past. I'm not going to comment on much from OP and instead give a little of my experience there. Straight out of college in 2017 I joined the Compute Fabric Controller (FC) org as a SWE on an absolutely wonderful team that dealt with mostly container management, VM and Host fault handling & repair policies, and Fabric to Host communication with most of our code in the FC. I drove our te…

> There was high turnover from the lack of headcount and overwork which was somewhat alleviated by lowering the hiring bar...

Seen this game played before, at AWS working on the control plane for outposts. The correct solution here is dedicated operations staff to coordinate with the team and let the developers fast track issues that are resulting in high call volumes, not lowering the hiring bar for the entire team. The problem you run into with high call volumes and small teams is that it disrupts most developers enough that they can't build solutions and deal with the maintenance burden at the same time. You bleed talent because it places way more stress than necessary on the team.

Re: Decisions that eroded trust in Azure – by a former Azure Core engineer

#622

Earlier quoted context omitted.

The first and most important lesson, that I try to each every young developer starting in the industry: Go home after clocking in your hours negotiated in your contract. Drop your pen. Go home. Sleep well. And I hope, that every sensible senior developer in here does the same. Lead by example. Maybe it would prevent a few burnouts in this industry. And if you are a manager, then send your people home after they have…

Good luck with that when you’re oncall.

On call should go against negotiated hours like 1/3 or 1/4. 3-4 8-hour shifts on call = day off. A single shift requiring active firefighting = day off.

Re: Decisions that eroded trust in Azure – by a former Azure Core engineer

#623

Earlier quoted context omitted.

The rapid decay of WTF/day over time applies to both new employees and new customers. > currently working on a legacy system "Legacy" is the magic word here! Those customers are pissed , trust me, but they've long ago given up trying to do anything about it. That's why you don't hear about it. Not because there are no bugs, but because nobody can be bothered to submit bug reports after learning long ago that doing so…

The customers aren't pissed, we're doing demos to new departments and lining up customizations and expansion as quickly as we can. We're growing faster than ever within our largest customer. I also didn't say there are no bugs or complaints, I said the system is more stable. But yes, there are fewer bugs and complaints, especially on the critical features. I didn't use the word legacy to mean abandoned, just that it'…

> But yes, there are fewer bugs and complaints

How do you know?

By that question I mean: Do you think there are fewer bugs because you hear fewer complaints from humans, or because you have a no-humans-involved mechanism for objectively evaluating the rate of bugs?

Even if you have a mechanical method for collecting bug reports, crash logs, or whatever, that can still obscure the true quality of the codebase.

One such example that I keep thinking about was the computer game Path of Exile. It has "super fans" that all have 10,000 hours of playtime that will swear up and down that it is one of the best games ever. When I first played it, I found so many little bugs and issues that I had more fun jotting them down than actually playing the game! I collected pages and pages of bullet points. None were crash bugs that would have been logged, and every one was the type of thing that players would eventually learn to work around by avoiding scenarios that caused the issue. I.e.: "Don't click to fast after going through a door because your orientation will be random on the other side, so you might be sent back to where you came from", that kind of thing.

Honestly and objectively measuring the quality of a software application (or any product) is hard.

Re: Decisions that eroded trust in Azure – by a former Azure Core engineer

#624

Earlier quoted context omitted.

> From another former Az eng now elsewhere still working on big systems, the post gets way way more boring when you realize that things like "Principle Group Manager" is just an M2 and Principal in general is L6 (maybe even L5) Google equivalent. Similarly Sev2 is hardly notable for anyone actually working on the foundational infra. Before the days of title inflation across the industry, a a Principal at Microsoft wa…

One of Microsoft's problems is their pay is significantly lower than FAANG and so you very very rarely see people with expertise in the same verticals jump to Azure. I get that "the deal" at Microsoft is lower pressure for lower pay but it really hinders the talent pipeline. There are some good home grown principals and seniors, but even then I think the people I worked with would have done well to jump around and ge…

> I get that "the deal" at Microsoft is lower pressure for lower pay but it really hinders the talent pipeline.

The deal used to be a lower cost of living in a major coastal city, an amazing campus (it is seriously lovely), every engineer had their own office, serious job security, and an unbelievable health care plan.

Seattle exploded in price, they moved to open offices, Microsoft started doing mass layoffs, and they gutted the healthcare plan (by the time I left the main plan on offer was a high deductible with a miserable prescription formulary).

Hard to attract talent when there is no big differentiator.

Of course in the 90s the deal was work there 10 years retire a millionaire. Easy to attract talent when that is the offer ...

Re: Decisions that eroded trust in Azure – by a former Azure Core engineer

#625

Earlier quoted context omitted.

I was once in such a position. I persuaded management to first cover the entire project with extensive test suite before touching anything. It took us around 3 months to have "good" coverage and then we started refactor of parts that were 100% covered. 5 months in the shareholders got impatient and demanded "results". We were not ready yet and in their mind we were doing nothing. No amount of explanation helped and t…

I’m a developer and if a team spent five months only refactoring with zero features added I would fire you too. Refactoring and quality improvements must happen incrementally and in parallel with shipping new features and fixing bugs.

I'm a director and one of our teams just spent 8 months doing just that and it was totally justified. They're finally coming up for air and the foundation is significantly improved.

There's nuance here. Every project/team/org is different.

Re: Decisions that eroded trust in Azure – by a former Azure Core engineer

#627
post #610

Earlier quoted context omitted.

Yeah, uhh: > I've worked on honing my communication skills for 20 years in this industry. That's because the skills weren't good enough.

So the takeaway isn't how good or bad I may be at communicating, it's that I was fundamentally speaking a language that was wholly orthogonal to the interests of leadership. No matter how good I became at making persuasive arguments about fixing technical debt and preventing outages, the management simply didn't care about those things. They say they they do, because it would sound insane to say otherwise, but they l…

Which is why you both listen to what they say, and pay attention to what they do, and what they prioritise. You use the actions to figure out where they were coming from with the message, and then you adapt your message to suit that.

> Which for many engineers who got into this industry because they loved solving problems, it can be quite a shocking realization.

It's just another problem to solve, based on the same foundational skill set you develop as an engineer: Observation, interpretation, analysis, experimentation, and implementation.

All-hands meetings are boring as hell, but they'll give me all sorts of signal about various managers up the line. I'll also take any opportunity I can get to be "in the room where it happens" when decisions are made (or speak to people who were in the room) while I'm building up a mental picture of what motivates someone.

If they're glory hunters, I'll figure out how to pitch my thing as something they can brag about. If they're people oriented (rare, but it happens), I'll pitch the human impact angle. If they're money pinchers, it's all about that $/month savings figure, put it front and centre in the opening sentence.

Everyone has an angle, a bias of some description. If you watch what projects do and don't get approved, and what language was used in them, you'll be successful too.

Re: Decisions that eroded trust in Azure – by a former Azure Core engineer

#628
I have always wanted to find some technical refutes, and I found one on reddit.

https://www.reddit.com/r/programming/comments/1sbir8j/commen...

I'll skip the other comments and focus on the technical ones:

> there are hundreds of “agents” which run on a one time basis to install systems as part of deployment architecture. These agents often amount to pretty simple scripts or programs. They most often run one time per update deployment, or if nodes are repaved. Some install small daemons. It’s called micro service architecture. Guy claims to be some cloud wiz but doesn’t get these basics.

> That said He’s put cutlers original work on a pedestal, when fabric controller should have been replaced a decade ago. The monolithic nature of fabric has been a huge issue for reliability and scalability, and the company is trying its hardest to move as many features out of it into microservices as it can.

I'm wondering if OP can answer this refute? Looks like the person is working in a neighbor team. No offense intended but I'm really curious about the technical part.

Re: Decisions that eroded trust in Azure – by a former Azure Core engineer

#629
The problem started because Azure was initially designed and released in a huge rush because Microsoft was so far behind AWS and needed something better.

I am reminded of the research finding that every human-designed complex system that works well started with a simple system that did did just one thing well, and new functions were added one at a time, with each one perfected before moving on to another. Which is the exact opposite of what happened here.

Re: Decisions that eroded trust in Azure – by a former Azure Core engineer

#630
post #249

Earlier quoted context omitted.

Once you reach this stage, the only escape is to jump ship. Either mentally or, ideally, truly. You're in an unwinnable position. Don't take the brunt for management's mistakes. Don't try to fix what you have no agency over.

unfortunately, what you will find is that unless you get lucky, the next ship is more of the same. The system/management style is ingrained in corporate culture of large-ish companies (i would say if it has more than 2 layers of management from you to someone owning the equity of the business and calling the shots, it's "large"). It stems from the fact that when an executive is bestowed the responsibility of managing…

> I would say if it has more than 2 layers of management from you to someone owning the equity of the business and calling the shots, it's "large"

By that metric, my 50 employee company is "large".

Post reply on HN