Live data from Hacker News

Ask HN: Do you find working on large distributed systems exhausting?

news.ycombinator.com

191–200 of 261 posts

Re: Ask HN: Do you find working on large distributed systems exhausting?

#191
post #37

I've found that external tech requirements are horrible to work with, especially when the underlying stack simply doesn't support it. Normally these are pushed by certified cloud consultants or by an intrepid architect who found another "best practice blog." It's begins with small requirements such as coming up with a disaster recovery plan only for it to be rejected because your stack must "automatically heal" and d…

I think this highlights the importance of actually analyzing your RP/RT (recovery point/recovery time) requirements through the lens of business value, and being honest about the ROI of buying that extra 9 in uptime.

It may be the case that 2 hours of downtime is completely unacceptable for the business, and paying $Xmm extra per year to maintain it is the right call. Or it may be that the business would be horrified to learn how many dollars are being spent to avert a level of downtime that no customer would notice or care about.

If the requirement is just being set by engineering, then it's more about finding the equilibrium where the resource spent on automation balances the cost of the manual toil and the associated morale impact on the team. Nobody wants to work on a team where everything is on fire all the time, and it's time/money well spent to avert that situation.

Re: Ask HN: Do you find working on large distributed systems exhausting?

#192
post #158

Earlier quoted context omitted.

You could have used a sharded database like Mongo. Just throw up 10 shards, use "source" (influencer or brand name) as shard key?

Yes, I could have used Mongo, but it would have been 100x to 1000x slower than an mmap-ed look up table.

Why ever use mmap instead of sharded inverted indices of word-doc here, a la elasticsearch?

Re: Ask HN: Do you find working on large distributed systems exhausting?

#194
post #173

Earlier quoted context omitted.

> Yeah, it was a dysfunctional environment and I obviously quit What do you think could management have done better to make it not dysfunctional and have people quitting?

I think just common sense and less bullshit rationalisation would have been enough. They had a billion dollars in cash to burn, so they hired more than they needed. They should have hired as needed, not as requested by Masayoshi Son. They shouldn't be so dogmatic. Some teams were too overworked, most were underworked (which means over-engineering will ensue), but no mobility was allowed because "ideally teams have N…

Favorited (https://news.ycombinator.com/favorites?id=akkartik&comments=...)

Re: Ask HN: Do you find working on large distributed systems exhausting?

#195

Earlier quoted context omitted.

Pre covid I would have laughed at this. But now, no one knows what a user story should be unless you can reas it off jira and there are no backups of course.

Gives me a fun idea: a program that randomly deletes items out of your backlog.

"Chaos engineering for your backlog"

Re: Ask HN: Do you find working on large distributed systems exhausting?

#196
post #177

Earlier quoted context omitted.

GP said they have never work on something that truly needed 50+ engineers. Truly being the keyword here IMO. I have worked on a 1000+ engineer project and another that was 500+, but I'm on the same boat as GP. Both of those didn't needed 50+, and the presence of the extra 950/450 caused several communication, organisational and architectural issues that became impossible to fix on the long term. So I can definitely s…

I've long wondered what I might be able to keep an eye out for during onboarding/transfer that would help me tell overstuffed kitchens apart from optimally-calibrated engineering caves from a distance. I'm also admittedly extremely curious what (broadly) had 1000 (and 500) engineers dedicated to it, when arguably only 50 were needed. Abstractly speaking that sounds a lot like coordinational/planning micromanagement,…

> a lot like coordinational/planning micromanagement, where the manglement had final say on how much effort needed to be expended where instead of allowing engineering to own the resource allocation process

Yep, that's a fair assessment!

The 1000+ one was an ERP for mid-large businesses. They had 10 or so flagship products (all acquired) and wanted to consolidate it all into a single one. The failure was more on trying to join the 10 teams together (and including lots of field-only implementation consultants in the bunch), rather than picking a solid foundation that they already owned and handpicking what needed.

The 500+ was an online marketplace. They had that many people because that was a condition imposed by investors. People ended up owning parts of a screen, so something that was a "two-man in a sprint" ended up being a whole team. It was demoralising but I still like the company.

I don't think it's impossible to notice, but it's hard... you can ask during interviews about numbers of employees, what each one does, ask for examples of what each team does on a daily basis. Honestly 100, 500, 1000 people for a company is not really a lot, but 100, 500, 1000 for a single project is definitely a red flag for me now, and anyone trying to pull the "but think of the scale!!!" card is a bullshit artist.

Re: Ask HN: Do you find working on large distributed systems exhausting?

#197
post #163

Earlier quoted context omitted.

So you basically saying that no matter what one should always stick to Amazon. I have my own experience that tells exactly the opposite. To each their own. We do not have to agree.

>So you basically saying that no matter what one should always stick to Amazon. What I am saying is given the list of exceptions I gave the business should run/colocate their gear if they're in the exception list or those components that fall in the exception list should be moved out. >I have my own experience that tells exactly the opposite. You begin using AWS for your first day ever and on that day it has a tri AZ…

>"You begin using AWS for your first day ever"

I am not talking about outages here. Bad things can happen. More like a price.

Re: Ask HN: Do you find working on large distributed systems exhausting?

#198
I think it's more likely Zeitgeist. You see, someone else finds working in Data Science frustrating, another person nearing his 40 says he's anxious about his career, another guy says he's worried about it's too late to do something about the big tech messing up the field etc.

I've had similar issues recently working at a demanding position I didn't really like even though my achievements may look impressive in my resume. I tried working in a shop somewhere in between aerospace and academia but just didn't fit at all. I ended up joining a small team that I enjoy working with so far and feel much better now.

At a higher level, we're hitting the limits of current paradigm in many ways including monetary system (debt), environment (pollution) and natural resources, ideology (creativity and innovation), technology (complexity).

The good news is that this year current monetary system will cease to exist. This will eventually restructure the economy to a more healthy balance. Unfortunately, this will have severe social consequences as standard of living will change dramatically (somewhere at the 60's level). This will basically destroy the middle class and thus change the structure of consumption. Obviously, this will mostly affect services and other non-essential stuff we got used to. On the other hand, this will blow down all bloat like insane market cap of the big tech etc. That is working in IT may become fun again, like 20 years back :)

Re: Ask HN: Do you find working on large distributed systems exhausting?

#200
post #9

What do you find exhausting? One anti-pattern I've found is that most orgs ask a single team to handle on-call around the clock for their service. This rarely scales well, from a human standpoint. If you're getting paged at 2:00 in the morning on a regular basis you will start to resent it. There's not much you can do about that so long as only one team is responsible for uptime 24/7. The solution is to hire operatio…

I would respectfully say that you are wrong. I speak from experience. At Netflix we tried to hire for around the clock coverage. But what ended up working much better was taking that same team and having each person on call for a week at a time, all based in Pacific Time. Yes, you would get calls at 2am, sometimes multiple days in a row. But you were only on call once every six to eight weeks, and we scheduled out we…

Even as a dedicated operations team for a product, we did this too. On call person worked tickets and took calls for one week at a time, the rest of the team worked on ways to make on-call suck less. For an eight person team it worked well for about three years until bigger stuff happened in the org and we all parted ways.
Post reply on HN