Live data from Hacker News

Ask HN: Do you find working on large distributed systems exhausting?

news.ycombinator.com

221–230 of 261 posts

Re: Ask HN: Do you find working on large distributed systems exhausting?

#221
post #220

Earlier quoted context omitted.

> Ah, so you worked on a team where the SRE needs were prioritized over the feature requests? Yes, it was an SRE team. All we do is write tools to make operations better, but more importantly we write tools to make it easier for the dev teams to operate their own systems better. But yes, we had products teams that would push back on our requests because they had product to deliver, and that was fine. We'd either figu…

> Step one, double the limit to alleviate immediate customer pain. I've been oncall for systems where that would not work. Doubling the memory means you need twice as many machines. Depending on the service, that could require significantly increased network bandwidth. Now the network is saturated and every node needs to queue more data. Now latency and throughput are even worse, and even more requests are being drop…

While that all may be true (but are indications of a poorly architected system), my code would still work. It would double the limit and then page someone. If they logged in and saw all those failures, then they could address those issues.

The whole point is that having an around the world follow the sun team would not alleviate those issues or make anything better.

Re: Ask HN: Do you find working on large distributed systems exhausting?

#222

Earlier quoted context omitted.

My impression has always been that FAANG need lots of engineers because the 10xers refuse to work there. I've seen plenty of really scalable systems being built by a small core team of people who know what they are doing. FAANG instead seem to be more into chasing trends, inventing new frameworks, rewriting to another more hip language, etc. I would have no idea how to coordinate 200 engineers. But then again, I have…

Your impression comes from the fact that you have not worked at larger teams, as you said so yourself. It's relatively easy to build something scalable from the beginning if you know what you need to build and if you are not already handling large amounts of traffic and data. It's a whole different ballgame to build on top of an existing complex system already in production that was made to satisfy the needs at the t…

Yeah, exactly. There is overhead simply because of the (necessary) cross-communication at that scale, and there's overhead from legacy support, but here's a thought experiment. Imagine that you've built the most perfect system from scratch that you can think of. Fast forward five years, and the business has pivoted so many times that system is doing all sorts of stuff it just wasn't designed for, and it's creaky and old. It just doesn't fit right anymore and even you want to throw it away and build a new one. So you form a tiger team full of the smartest people you know to greenfield build a new one, from scratch, but that's gonna take two years to write. (You think, hey, maybe we could just take this open source thing and adapt it to our purposes. To which I say, where do you think large open source projects come from‽)

How do you bridge the two systems? You build an interim system. But customers want new features, so those features need to be done twice (bridge+new) if you're lucky, three times (existing+interim+new) if not. Could a smaller team of 10x engineers come in and do better? First off, thanks for insulting all of us, as if none of us are 10x-ers. But no. There's simply not enough hours in the day.

We've all heard of large IT projects that failed to land and said "of course". But we don't hear about the huge ones that do. And plenty of them do land, quite succesfully, with these 200+ person teams where I, as an SRE, don't know the code for the system I'm supporting.

None of this is visible from the outside.

Re: Ask HN: Do you find working on large distributed systems exhausting?

#223
Is it really the distributed aspect? Or "just" working on a above average complicated project for many years?

The consequences of bugs in many distributed systems (and several other types of systems) are IME often harder to bear than e.g. UI or frontend workflow bugs. It's hard to have caused data loss. And at some point you probably will, even if you're quite careful.

Maybe I'm just projecting...

Re: Ask HN: Do you find working on large distributed systems exhausting?

#224
post #173

Earlier quoted context omitted.

> Yeah, it was a dysfunctional environment and I obviously quit What do you think could management have done better to make it not dysfunctional and have people quitting?

I think just common sense and less bullshit rationalisation would have been enough. They had a billion dollars in cash to burn, so they hired more than they needed. They should have hired as needed, not as requested by Masayoshi Son. They shouldn't be so dogmatic. Some teams were too overworked, most were underworked (which means over-engineering will ensue), but no mobility was allowed because "ideally teams have N…

> They shouldn't be so dogmatic pt 2. Services were one-per-team, instead of one-per-subject.

Where the heck did this come from? AIUI, the ideal is supposed to be one-team-per-service, not one-service-per-team.

Re: Ask HN: Do you find working on large distributed systems exhausting?

#225
post #173

Earlier quoted context omitted.

I think just common sense and less bullshit rationalisation would have been enough. They had a billion dollars in cash to burn, so they hired more than they needed. They should have hired as needed, not as requested by Masayoshi Son. They shouldn't be so dogmatic. Some teams were too overworked, most were underworked (which means over-engineering will ensue), but no mobility was allowed because "ideally teams have N…

> They shouldn't be so dogmatic pt 2. Services were one-per-team, instead of one-per-subject. Where the heck did this come from? AIUI, the ideal is supposed to be one-team-per-service, not one-service-per-team.

It comes from a dogmatic reaction against microservices. Microservices were problematic in certain ways, but instead of analysing what went wrong and why, they just went the opposite direction and started doing "big services only". It was a misguided approach, plain and simple.

Interestingly due to internal bureaucracy and understaffing in some teams, there was a lot of "multiple-teams-per-service", which yeah, is another issue in itself.

Re: Ask HN: Do you find working on large distributed systems exhausting?

#226
post #65

Earlier quoted context omitted.

I feel that your problems aren't even remotely related to my problems with large distributed systems. My problems are all about convincing the company that I need 200 engineers to work on extremely large software projects before we hit a scalability wall. That wall might be 2 years in the future so usually it is next to impossible to convince anyone to take engineers out of product development. Even more so because w…

My impression has always been that FAANG need lots of engineers because the 10xers refuse to work there. I've seen plenty of really scalable systems being built by a small core team of people who know what they are doing. FAANG instead seem to be more into chasing trends, inventing new frameworks, rewriting to another more hip language, etc. I would have no idea how to coordinate 200 engineers. But then again, I have…

Which FAANG is rewriting to another hip language and chasing trends (especially when it comes to infra services??)? I don't mean to be rude, but it doesn't sound like you are talking about any of the FAANGs, this sounds completely made up.

Re: Ask HN: Do you find working on large distributed systems exhausting?

#227

My experience is that the expectations on what your average engineer should be able to handle has grown enormously during the last 10 years or so. Working both with large distributed systems and medium size monolithic systems I have seen the expectations become a lot higher in both. When I started my career the engineers at our company were assigned a very specific part of the product that they were experts on. Usual…

Might I add that you are also now underpaid. I had a sweet gig at a very small company where I had to manage contractors in addition to FTE staff. The good contractors billed $300 an hour for BA and project management services alone. The story munchers billed $150 an hour.

I had to leave a contracting gig recently because we were tasked with everything...literally everything. Everyone got so burnt out--FTEs included. I also might add that the developers could have spoken up and gotten relief but their misguided work ethic prevented that.

Re: Ask HN: Do you find working on large distributed systems exhausting?

#228
post #35

The first ten years of my career, I worked with distributed systems built on this stack: C++, Oracle, Unix (and to some extent, MFC and Qt). There were hundreds of instances of dozens of different type of processes (we would now call these microservices) connected via TCP links, running on hundreds of servers. I seldom found this exhausting. The second ten years of my career, I worked with (and continue to work on) m…

Modern day embarrassing spaghetti cloud.

Re: Ask HN: Do you find working on large distributed systems exhausting?

#229
post #147

Earlier quoted context omitted.

I am not sure you are aware that server load is never linearly distributed. And that's the exact problem OP is talking about. If everybody would get a ticket number and do requests when they're supposed to do them, we wouldn't need load balancers.

This is orthogonal to what causes the pain point. All the pain comes from distributed state and these load levels even if you peak at 800K requests per second you don't need distributed state. So most of this pain is self inflicted.

True, I agree.

In most systems I've seen, the caching layer is invalidated more often than necessary, and most of the traffic could've been avoided with a better URL scheme that's more expressive in regards to its content (mutations).

Re: Ask HN: Do you find working on large distributed systems exhausting?

#230
post #198

I think it's more likely Zeitgeist. You see, someone else finds working in Data Science frustrating, another person nearing his 40 says he's anxious about his career, another guy says he's worried about it's too late to do something about the big tech messing up the field etc. I've had similar issues recently working at a demanding position I didn't really like even though my achievements may look impressive in my re…

The end is near?

Looks like that. I mean has anything fundamentally changed since 2008? No. The 'raegonomical' approach has been eating the future demand for 40 years. And what is the dollar issue volume in 2020-2021 compared to preexisting volume? And what's been happening to PPI[1] since 2020? I guess we should ask the Fed about that.

[1] https://fondmx.pro/wp-content/uploads/2022/02/image-174.png

Post reply on HN