My hope is that I'll learn to manage the stress and gain more expertise.
Ask HN: Do you find working on large distributed systems exhausting?
171–180 of 261 posts
Re: Ask HN: Do you find working on large distributed systems exhausting?
#172Re: Ask HN: Do you find working on large distributed systems exhausting?
#173Earlier quoted context omitted.
> Why do large sites like Facebook, Amazon, Twitter and Instagram all essentially look the same after 10 years but some of them now have 10x the amount of engineers? I think they have so much data and so many dependencies between parts of the system that any fundamental change is extremely hard to pull off. They even cut back on features like API access. But I am pretty sure that most of them have rewritten the whole…
> Yeah, it was a dysfunctional environment and I obviously quit What do you think could management have done better to make it not dysfunctional and have people quitting?
They had a billion dollars in cash to burn, so they hired more than they needed. They should have hired as needed, not as requested by Masayoshi Son.
They shouldn't be so dogmatic. Some teams were too overworked, most were underworked (which means over-engineering will ensue), but no mobility was allowed because "ideally teams have N people".
They shouldn't be so dogmatic pt 2. Services were one-per-team, instead of one-per-subject. So yeah, our internal tool for putting balloons and clowns into images lived together with the authentication micro-service, because it's the same team.
Rewriting everything twice without analysis was wrong. The rewrites were because previous versions were "too complex" and too custom-made but newer ones had an even more complex architecture, but "this time it's right, software sometimes need complexity".
Believing that some things were terrible would have gone a long way. Launching the main node.js server locally would take 10 to 20 minutes to launch, while something of the same complexity would often take about 2 or 3 seconds. Of course it would blow up in production! Maybe try to fix instead of ordering another rewrite.
They were good people, I miss the company and still use the product, but it didn't need to be like this.
Re: Ask HN: Do you find working on large distributed systems exhausting?
#174Earlier quoted context omitted.
>This sounds a bit arrogant. The parent thread talks about how the business could not go down even with a triple AZ outage for S3, and I don't think it is arrogant to state they should be paying for enterprise support if that level of expectation is set. >I think they found better and overall cheaper solution. Cheaper solution does not just include the cost but also the time. For the time we need to look at the time…
So you basically saying that no matter what one should always stick to Amazon. I have my own experience that tells exactly the opposite. To each their own. We do not have to agree.
What I am saying is given the list of exceptions I gave the business should run/colocate their gear if they're in the exception list or those components that fall in the exception list should be moved out.
>I have my own experience that tells exactly the opposite.
You begin using AWS for your first day ever and on that day it has a tri AZ outage for S3. In this example the experience with AWS has been terrible. Zooming out though over 5 years it wouldn't look like a terrible experience at all considering outages are limited and honestly not that frequent.
Re: Ask HN: Do you find working on large distributed systems exhausting?
#175The most undervalued thing that forgot even highly skilled engineers - KISS principle. That’s why you are burning out supporting such systems.
Yes, it's amazing how much one modern high spec system running good code can do. Turn off all the distributed crap and just use a pair in leader/follower config with short ttl DNS to choose the leader and manual failover scripts. If your app/company/industry cannot accept the compromises from such a simple config, quit and work in one which can.
This whole thread feels like therapy since I face the same monsters on the systems I work on. Partly due to bad platform & code, partly due to bad organization structure (Conway's 100% for us).
My pet projects at home is the only thing keeping me sane, mostly because they are simple.
Re: Ask HN: Do you find working on large distributed systems exhausting?
#176Earlier quoted context omitted.
Typically people use RAID or ZFS to prevent data loss when a few hdds fail.
OK, so basically you're in a completely different class of expectations about how systems perform under disk loss and heavy load then me. A drive array is very different from large-scale cloud storage.
- A large ZFS pool of SSDs is much faster than any cloud storage.
- Cloud storage failed much more often than the SSDs in our pool.
- "Noisy neighbor" is an issue on the cloud
Re: Ask HN: Do you find working on large distributed systems exhausting?
#177Earlier quoted context omitted.
My impression has always been that FAANG need lots of engineers because the 10xers refuse to work there. I've seen plenty of really scalable systems being built by a small core team of people who know what they are doing. FAANG instead seem to be more into chasing trends, inventing new frameworks, rewriting to another more hip language, etc. I would have no idea how to coordinate 200 engineers. But then again, I have…
Your impression comes from the fact that you have not worked at larger teams, as you said so yourself. It's relatively easy to build something scalable from the beginning if you know what you need to build and if you are not already handling large amounts of traffic and data. It's a whole different ballgame to build on top of an existing complex system already in production that was made to satisfy the needs at the t…
I have worked on a 1000+ engineer project and another that was 500+, but I'm on the same boat as GP. Both of those didn't needed 50+, and the presence of the extra 950/450 caused several communication, organisational and architectural issues that became impossible to fix on the long term.
So I can definitely see where they're coming from.
Re: Ask HN: Do you find working on large distributed systems exhausting?
#178Earlier quoted context omitted.
>This sounds a bit arrogant. The parent thread talks about how the business could not go down even with a triple AZ outage for S3, and I don't think it is arrogant to state they should be paying for enterprise support if that level of expectation is set. >I think they found better and overall cheaper solution. Cheaper solution does not just include the cost but also the time. For the time we need to look at the time…
Back then, it was enough to saturate the S3 metadata node for your bucket and then all AZs would be unable to service GET requests. And yes, this won't be financially useful in every situation. But if the goal is to gain operational control, it's worthwhile nonetheless. That said, for a high-traffic API, you're paying through the nose for AWS egress bandwidth, so it is one of those cases where it also very much makes…
Re: Ask HN: Do you find working on large distributed systems exhausting?
#179Yes, I used to, but No, I fixed it :) Among other things, I am team lead for a private search engine whose partner-accessible API handles roughly 500 mio requests per month. I used to feel powerless and stressed out by the complexity and the scale, because whenever stuff broke (and it always does at this scale), I had to start playing politics, asking for favors, or threatening people on the phone to get it fixed. Hi…
I feel that your problems aren't even remotely related to my problems with large distributed systems. My problems are all about convincing the company that I need 200 engineers to work on extremely large software projects before we hit a scalability wall. That wall might be 2 years in the future so usually it is next to impossible to convince anyone to take engineers out of product development. Even more so because w…
sounds like a typical massive rewrite project. They almost never succeed, many fail outright and most hardly even reach the functionality/performance/etc. level of the stuff the rewrite was supposed to replace. 2-4 years is typical for such glorious attempt before being closed or folded into something else. Management in general likes such projects, and they declare victory usually around 2 years mark and move on on the wave of the supposed success before reality hits the fan.
>to convince anyone to take engineers out of product development.
that means raiding someone's budget. Not happening :) New glorious effort needs new glorious budget - that is what management likes and not doing much more on the same budget as you're basically suggesting (i.e. i'm sure you'll get much more traction if you restate your proposal as "to hire 200 more engineers ..." because that way you'll be putting serious technical foundation for some mid-managers to grow :). You're approaching this as an engineer and thus failing in what is the management game (or as Sun Tzu was pointing out one has to understand the enemy).
Re: Ask HN: Do you find working on large distributed systems exhausting?
#180Recently I was asked to work on a older project for enterprise customers. And we are always weary of working on old unmaintained code But it just felt like a breath of fresh air All code in same repository, UI, back-end, SQL, MVC style Fast from feature request to deliver in production. Changes, test, fix bugs, deploy. We were happy and the customers too No cloud apps, buckets, secrets, no oauth, little configuration…