Live data from Hacker News

Ask HN: Do you find working on large distributed systems exhausting?

news.ycombinator.com

231–240 of 261 posts

Re: Ask HN: Do you find working on large distributed systems exhausting?

#231

Earlier quoted context omitted.

> you were highly incentivized (and had the time) to build tools or submit patches to fix the problems that woke you at 2am. Ah, so you worked on a team where the SRE needs were prioritized over the feature requests? Because in most companies where I've worked, Product + Customer Service + Sales + Marketing + Executives don't really have time or patience for the engineers to get their diamond polishing cloths out. Th…

> Ah, so you worked on a team where the SRE needs were prioritized over the feature requests? Yes, it was an SRE team. All we do is write tools to make operations better, but more importantly we write tools to make it easier for the dev teams to operate their own systems better. But yes, we had products teams that would push back on our requests because they had product to deliver, and that was fine. We'd either figu…

> Yes, it was an SRE team. All we do is write tools to make operations better

Go back to my original comment. If you're an SRE team, then basically, you're the operations team for the developers. I'm talking about where developers are responsible for their own operations and there is no team that gets paged instead of them - "most orgs ask a single team to handle on-call around the clock for their service."

> Step one, double the limit to alleviate immediate customer pain. Step two, page someone to wake up

See, what I read from this is: a) violate my system efficiency KPIs while b) paging someone in the middle of the night anyway. So, lose-lose.

> But it makes no sense to hire an around the world team just for operations if the rest of your company is in one time zone.

Why does it make any more sense to hire developers remotely who are in your time zone ± three hours? Because that's what most companies are doing these days. If you're already hiring people remotely then you can hire Operations/SRE staff a little further afield and see that as a benefit (follow the sun) rather than a problem.

> the rest of your company is in one time zone.

For what it's worth, we also hire salespeople around the globe :) Fact of the matter is, it would be so, so nice for Slack to turn off the ability to @channel in the #random channel so that people who are asleep don't get pinged ...

Re: Ask HN: Do you find working on large distributed systems exhausting?

#232

Earlier quoted context omitted.

> Ah, so you worked on a team where the SRE needs were prioritized over the feature requests? Yes, it was an SRE team. All we do is write tools to make operations better, but more importantly we write tools to make it easier for the dev teams to operate their own systems better. But yes, we had products teams that would push back on our requests because they had product to deliver, and that was fine. We'd either figu…

> Yes, it was an SRE team. All we do is write tools to make operations better Go back to my original comment. If you're an SRE team, then basically, you're the operations team for the developers. I'm talking about where developers are responsible for their own operations and there is no team that gets paged instead of them - "most orgs ask a single team to handle on-call around the clock for their service." > Step on…

We were an SRE team building tools for the development teams who got paged in the middle of the night. The devs writing the services were operating their own services and were getting paged. We would sometimes also get paged for a serious incident so we could coordinate if multiple development teams were involved.

Each team managed their own rotation schedules, we just made sure they had one.

> See, what I read from this is: a) violate my system efficiency KPI

If you're being graded on your system efficiency and not customer satisfaction, well then sure, your way might make sense (but I'd still say it doesn't). But your business will suffer if you optimize for efficiency over customer satisfaction.

> Why does it make any more sense to hire developers remotely who are in your time zone ± three hours?

Because it's a lot easier to run a team where everyone on the team can meet at the same time. If you have an around the world team, there is no time of day where you can have a meeting and everyone gets to attend during their workday. Realistically you can maybe get away with a nine hour time difference. Any more than that and you have people excluded.

Especially if the bulk of your devs are in one or two time zones, your operators will be even more disconnected from them since they will never be able to interact with the devs, and the devs will have no empathy for the operators who they also never interact with.

> For what it's worth, we also hire salespeople around the globe

Sure, but they aren't writing code that your operators have to run. :)

I think we both agree that it's better for devs to get paged for their services instead of operators, and if that's the case, its far better for all the devs to work together and know each other and be in the same or nearly same time zone.

A follow the sun model breaks that completely.

Re: Ask HN: Do you find working on large distributed systems exhausting?

#233

Earlier quoted context omitted.

> Yes, it was an SRE team. All we do is write tools to make operations better Go back to my original comment. If you're an SRE team, then basically, you're the operations team for the developers. I'm talking about where developers are responsible for their own operations and there is no team that gets paged instead of them - "most orgs ask a single team to handle on-call around the clock for their service." > Step on…

We were an SRE team building tools for the development teams who got paged in the middle of the night. The devs writing the services were operating their own services and were getting paged. We would sometimes also get paged for a serious incident so we could coordinate if multiple development teams were involved. Each team managed their own rotation schedules, we just made sure they had one. > See, what I read from…

> But your business will suffer if you optimize for efficiency over customer satisfaction.

But who are the customers? Business, engineering, or finance? :)

> it's a lot easier to run a team where everyone on the team can meet at the same time.

Of course it's easier. It's also easier not to maintain documentation or standards, just be a five person startup and have everyone be in the same room. Enterprise communication is hard! Even when you're in the same time zone. The question isn't "how do I get my life to be a utopia?" but "which challenges should I choose?". If you run an organization, you need to put your employees first, even ahead of your customers. Employees and customers both come and go but 80% of the time the effect of an valued employee leaving is far worse than a customer leaving, and you have far more control over whether employees leave than whether customers do. So you can either put your employees first (build a calm workplace) or you can put your customers first (prioritize feature development velocity in organizational design).

> I think we both agree that it's better for devs to get paged for their services instead of operators

No! Dev should never be paged! If I "buy" Jenkins off-the-shelf, and it breaks down in production, guess what, I don't get to page the Jenkins developers! Why should internally developed services be any different? If Ops needs to page someone from Dev instead of waiting for a response at normal business cadence, then this is an Ops failure, not a Dev failure!

Re: Ask HN: Do you find working on large distributed systems exhausting?

#234

Earlier quoted context omitted.

We were an SRE team building tools for the development teams who got paged in the middle of the night. The devs writing the services were operating their own services and were getting paged. We would sometimes also get paged for a serious incident so we could coordinate if multiple development teams were involved. Each team managed their own rotation schedules, we just made sure they had one. > See, what I read from…

> But your business will suffer if you optimize for efficiency over customer satisfaction. But who are the customers? Business, engineering, or finance? :) > it's a lot easier to run a team where everyone on the team can meet at the same time. Of course it's easier. It's also easier not to maintain documentation or standards, just be a five person startup and have everyone be in the same room. Enterprise communicatio…

> But who are the customers? Business, engineering, or finance? :)

The business's customers. The ones who pay your company so they can pay you, and your reason for having a job at all.

> Why should internally developed services be any different?

Because they're your core competency and you have control over it. If you could page the Jenkins developers you probably wouldn't hesitate to do it, because you'll get better results. Why not get the best results you can from an internal service?

> If Ops needs to page someone from Dev instead of waiting for a response at normal business cadence, then this is an Ops failure, not a Dev failure!

I couldn't disagree more. That is absolutely a dev failure -- they wrote a service that couldn't operate under the conditions given. It's either a bug or an architecture issue, but no matter what, it's a dev issue and the dev should be responsible for building a service that can actually run in production.

You and I have very different ideas of a successful engineering organization. I would never want to work for your org as an operator or a dev. As an operator the last thing I want is devs to throw whatever they write over the wall and then say "not my problem anymore!", and have to rely on getting retrained every time the code changes. And as a dev I wouldn't want to be in an organization that accepts sloppy developers who aren't responsible for building solid code that can run under adverse conditions and who don't get to experience the issues in production for themselves.

Facebook makes their devs get paged, Netflix does, Amazon pages their devs, Dropbox pages devs, Stripe pages devs, and Google pages their devs too until they have demonstrated multiple quarters of success, and only then does an operator take over. And if the service has too many failures, support falls back on the devs until they can make the service stable again.

Making devs responsible for creating code that actually works well in production is a good thing.

Re: Ask HN: Do you find working on large distributed systems exhausting?

#235

Earlier quoted context omitted.

> But your business will suffer if you optimize for efficiency over customer satisfaction. But who are the customers? Business, engineering, or finance? :) > it's a lot easier to run a team where everyone on the team can meet at the same time. Of course it's easier. It's also easier not to maintain documentation or standards, just be a five person startup and have everyone be in the same room. Enterprise communicatio…

> But who are the customers? Business, engineering, or finance? :) The business's customers. The ones who pay your company so they can pay you, and your reason for having a job at all. > Why should internally developed services be any different? Because they're your core competency and you have control over it. If you could page the Jenkins developers you probably wouldn't hesitate to do it, because you'll get better…

> As an operator the last thing I want is devs to throw whatever they write over the wall and then say "not my problem anymore!", and have to rely on getting retrained every time the code changes. And as a dev I wouldn't want to be in an organization that accepts sloppy developers who aren't responsible for building solid code that can run under adverse conditions and who don't get to experience the issues in production for themselves.

How can you classify anybody who writes on-prem software as being a "sloppy developer"? Jira, Jenkins, GitLab, pretty much any database you can imagine (MySQL, PostgreSQL, Redis, Elasticsearch, Kafka...), Grafana, any Linux distribution, they're all written by "sloppy developers"?

Where did I say that Dev gets to "throw code over the wall"? How would you feel if I unilaterally decided for you, as a developer, which tools you get to use? If I came up with some policy that the whole organization can only run Windows machines and I "threw that policy over the wall" at you?

You're arguing against a strawman that's completely inconsistent with how harmonious follow-the-sun Ops actually works.

Re: Ask HN: Do you find working on large distributed systems exhausting?

#236

Earlier quoted context omitted.

My impression has always been that FAANG need lots of engineers because the 10xers refuse to work there. I've seen plenty of really scalable systems being built by a small core team of people who know what they are doing. FAANG instead seem to be more into chasing trends, inventing new frameworks, rewriting to another more hip language, etc. I would have no idea how to coordinate 200 engineers. But then again, I have…

Which FAANG is rewriting to another hip language and chasing trends (especially when it comes to infra services??)? I don't mean to be rude, but it doesn't sound like you are talking about any of the FAANGs, this sounds completely made up.

https://blog.pragmaticengineer.com/uber-app-rewrite-yolo/

Re: Ask HN: Do you find working on large distributed systems exhausting?

#237
post #207

Earlier quoted context omitted.

you seem to have a deeper knowledge of the business & organisational context that dictate the true requirements than someone working there. please share these details so we can all learn!

Sure: the network request time of a person making a request over the open internet is going to be an order of magnitude longer than a DB lookup (in the right style, with a reverse-index) on the scale of data this person is describing. So making the lookup 10x faster saves you...1% of the request latency. And at the qps they've described, it's not a throughput issue either. So I'm pretty confident in saying that this…

This reads to me as if you have never really used mmap in a dedicated C/C++ application. Just to give you a data point, looking up one word_id in the LUT and reading 20 document_ids from it takes on average 0.0000015 ms.

So if that alternative database takes on average 0.1ms per index read, then it's starting out roughly 65000x slower.

"than a DB lookup (in the right style, with a reverse-index)"

Unless, of course, you're managing petabytes of data ;)

"at the qps they've described, it's not a throughput issue either"

It's mostly a cost thing. If a single request takes 2x the time, that's also a 2x on the hosting bill.

"parallelization of scans dominates mmap speed"

Yes, eventually that might happen. Roughly when you have 100000 servers. But before that your 10gbit/s node-to-node link will saturate. Oops.

Re: Ask HN: Do you find working on large distributed systems exhausting?

#238
One problem I frequently see with distributed systems is not the amount of services and the distributed nature per se.

Rather that it allows, and tempts, you to use the perfect tool for each job. Leading to a lot of variations in your stack.

Suddenly you have 5 different databases, 3 RPC protocols, 4 programming languages and 2 operating systems spinning around in your cluster. Only half of them connected to your single sign on. And don’t forget about all the cloud dependencies.

If any one of them starts misbehaving you have to read up “how did I attach debugger to Java process again”. How do I even log in to a mongodb shell? I installed pgadmin last week.

Standardize your stack and accept that some times it might mean using something slightly inefficient in the small scheme. In the big scheme it will make things more homogenous, unified and simpler for operators.

Re: Ask HN: Do you find working on large distributed systems exhausting?

#239
post #21

Yes, but in a different way. I work in Quality Engineering, and the scope of maturity in testing distributed systems has been exhausting. Reading other comments from the thread, I see similar frustrations from teams I partner with. How to employ patterns like contact, hypothesis, doubles, or shape/data systems (etc.) typically gets conflated with System testing. Teams often disagree on the boundaries of the system st…

I so wish there were in-person meetups and conferences going on so I might have been nearby and overheard you saying that so I could try to join in the conversation. Sounds fascinating and just the sort of insight that doesn't come up in the entirely planned and scheduled zooms I'm usually in (and HN, for all its virtues, isn't really a substitute for a great conversation).

Re: Ask HN: Do you find working on large distributed systems exhausting?

#240
post #9

What do you find exhausting? One anti-pattern I've found is that most orgs ask a single team to handle on-call around the clock for their service. This rarely scales well, from a human standpoint. If you're getting paged at 2:00 in the morning on a regular basis you will start to resent it. There's not much you can do about that so long as only one team is responsible for uptime 24/7. The solution is to hire operatio…

I would respectfully say that you are wrong. I speak from experience. At Netflix we tried to hire for around the clock coverage. But what ended up working much better was taking that same team and having each person on call for a week at a time, all based in Pacific Time. Yes, you would get calls at 2am, sometimes multiple days in a row. But you were only on call once every six to eight weeks, and we scheduled out we…

I agree with you completely, especially on the last paragraph. No pain - no gain.
Post reply on HN