Live data from Hacker News

Why Twitter didn’t go down: From a real Twitter SRE

matthewtejo.substack.com

911–920 of 1001 posts

Re: Why Twitter didn’t go down: From a real Twitter SRE

#911

Earlier quoted context omitted.

> In a start-up many of these roles can be ignored, becaus growth > stability. In a large organization, part of the bloat helps insure a certain amount of stability that's necessary to keep an organization alive. It also (a) increases the bus factor, [1] and (b) allows people to take vacations and time off without having to watch their phones like hawk. [1] https://en.wikipedia.org/wiki/Bus_factor

I know the bus factor is the morbid "how many people can get hit by a bus?" idea, but I actually like to present it as "how wide is the bus?" in terms of people being the conduits along which information and instructions flow; a wider bus provides redundancy. Aside from being more positive I think it's also more accurate for how we want teams to work. The "win the lottery and quit" metaphor is just stupid.

I don't think the analogy holds in the same way. To me, the bus factor represents how many people you can lose for an extended period of time before there are no subject matter experts left for a particular topic. If you've got 8 people on the team, but only 2 of them know how to do a particular thing, the bus factor is 2.

It's not about having enough people to do the work even if someone quits, it's about having enough people that know how to do something that we aren't losing chunks of knowledge if someone quits (or dies, or gets fired, or gets sick, or etc).

It doesn't make sense to me to treat people as part of a conduit bus that are interchangeable as long as there are enough people.

Re: Why Twitter didn’t go down: From a real Twitter SRE

#912

Earlier quoted context omitted.

I'm still on the fence. As an engineering manager, I tend to attach faces to those "jobs". So seeing cuts, I imagine a ton of people that had to go home and tell their friends/family/etc that they no longer had a job. On the other hand. As an engineer, we tend to attach way too much self importance to our roles. Like if we're not there entering the "numbers" 4 6 15 16 24 32 every 108 minutes, the entire business is g…

It's great to hear your concern with the actual people. That tends to get lost for some reason. And I agree we tend to attach too much importance to our roles, but the flip side bears a certain amount of truth, though on a timeline more like 108 days than 108 minutes.

I definitely walk a fine line of wanting people to know how much I value them and how valuable they are to the business as a whole. But, the harsh truth is, every single one of us is replaceable. I've just seen that play out too many times in my career. People with all kinds of domain tribal knowledge walk out the door. Everyone holds their breath as if it'll be the end. And a month or so later, we're still afloat. (I still hate to see good people leave)

Re: Why Twitter didn’t go down: From a real Twitter SRE

#913
post #27

All this does is point out that smart people worked at Twitter who may now no longer work there, whether on their own accord, or due to Elon’s bulldogging tactics. Elon thinks he knows what he’s doing, but what he is going to be left with are people who are willing to work hard by his standards, but not necessarily smart. The simple truth is Elon knows nothing about the actual work involved in tech. He knows words or…

Years ago I had a massively downvoted comment when I criticised his AR for CAD vapour ware. As someone who was fully in that area at the time, what he was showing while looking fancy had no practical application in the area of design he was talking about.

Ever watch someone do CAD/CAM modelling? They need extreme precision of input that AR sausage fingers just aren't going to help with. You need a num-pad and a good mouse with a stepped click wheel.

I get it. We share some similar neurodivergence traits. He wants to be right in the detail. Constantly jumping from interest to interest, seeing the hidden patterns and connections that aren't apparent to others. But there are a times when I know I just need to shut up and let someone more experienced talk despite my brain wanting to lead every discussion right into solution mode, or providing additional context mode.

I've spent the last 6 years in management consulting (without formal business education), I agree with him when he says MBAs are useless. We know that the best solutions come from diverse teams with diverse backgrounds, skills and knowledge. Not 5 clones who know how to build value driver trees, not to say the tools they bring aren't useful, but they can be incredibly limiting.

For someone who hates MBAs he's sure going about this take-over like someone who barely passed one (i.e. knows more than enough to be dangerous). Sure, you're hemorrhaging money in operations. You need to cut costs and find new revenue streams.

What are your biggest costs?

Labor. Slash / Burn. The old McKinsey 7% FTE reduction will give you some extra operating cash from the years remaining budget and you know it's not so much that people (in fear of their jobs) won't just pick up the slack to keep everything moving. Do it quick because you need to rip the band-aid off and get rid of all that accrued leave, restricted cash etc. off your books too.

Equipment. Redundancy? Sounds like unused resources we can fire sale.

Contracts. Renegotiate? The only two meaningful levers are price and quantity. Start cutting quantity now, renegotiate price later.

This is all dummies guide stuff and tends to go terribly in reality when implemented all at once all together.

For instance, research has shown companies that lay-off when under pressure end up underperforming against the ones who chose not to.

Now who's going to help build and operate those new revenue streams?

Quick fixes for a quick buck and a whole lot of extra risk.

Re: Why Twitter didn’t go down: From a real Twitter SRE

#914
post #609

> This left a lot wondering what exactly was going on with all those engineers and made it seem like it was all just bloat. I was partly expecting the rest of the article to explain to me why exactly it wasn't just bloat. But it goes on talking about this 1~3-person cache SRE team that built solid infra automation that's really resilient to both hardware and software failures. If anything, the article might actually…

You just need to look at profit and revenue. Revenue grew nicely in last 2-3 years, profit not so much. Bloat is only reason.

This, in my reflection, is that one insightful comment that should be higher up even it came to late. Twitter userbase did not expand significantly in the last couple of years. Revenue increased. Why did cost increase so much?

Re: Why Twitter didn’t go down: From a real Twitter SRE

#915
post #162

From this operations engineer's perspective, there are only 3 main things that bring a site down: new code, disk space, and 'outages'. If you don't push new code, your apps will be pretty stable. If you don't run out of disk space, your apps will keep running. And if your network/power/etc doesn't mysteriously disappear, your apps will keep running. And running, and running, and running. The biggest thing that brings…

Another thing we noticed at Netflix was that after services didn’t get pushed for a while (weeks), performance started degrading because of things like undiscovered memory leaks, threads leaks, disks filling up. You wouldn’t notice during normal operations because of regular autoscaling and code pushes, but code freezes tended to reveal these issues.

[deleted]

Re: Why Twitter didn’t go down: From a real Twitter SRE

#916

Earlier quoted context omitted.

These people behave just like the irrational fanboys, except they just do the exact opposite. Being a sheep and being a contrarian sheep are the same thing.

I've no dog in this race but I'm excited to see what influence, if any, this will have on the topology of similarly inflated tech companies.

Honestly, I'm hopeful that other coastal tech companies do clean house. I genuinely believe it could lead to a resurgence in the lives of Midwest/middle American states and cities.

There is an absolute vacuum of technology specialists in the middle of the US, because no one wants to "live in the middle of nowhere," and they don't want to earn less than FAANG (MAMAA?) salaries, when half of those salaries can give you an amazing life in the middle of the country (source: my piss-poor salary compared to yours).

Re: Why Twitter didn’t go down: From a real Twitter SRE

#917

Earlier quoted context omitted.

I know the bus factor is the morbid "how many people can get hit by a bus?" idea, but I actually like to present it as "how wide is the bus?" in terms of people being the conduits along which information and instructions flow; a wider bus provides redundancy. Aside from being more positive I think it's also more accurate for how we want teams to work. The "win the lottery and quit" metaphor is just stupid.

I usually rephrase it as “getting run over by the lottery” or “winning a bus”. Just to be different/humorous/less morbid.

Have I been in a life-long bubble such that the ordinary explanation of "bus factor" seems so mild that until reading exchanges like this a couple times it never occurred to me it might be morbid enough to bother someone? Or is this one of those things where people are looking for something to worry about but it's actually entirely fine? Like, I've seen children's cartoons with jokes that were more morbid than that.

(I do also like the lottery version and use them basically interchangeably, though the shorthand is always "bus factor" for me)

Re: Why Twitter didn’t go down: From a real Twitter SRE

#919
post #162

Earlier quoted context omitted.

Another thing we noticed at Netflix was that after services didn’t get pushed for a while (weeks), performance started degrading because of things like undiscovered memory leaks, threads leaks, disks filling up. You wouldn’t notice during normal operations because of regular autoscaling and code pushes, but code freezes tended to reveal these issues.

We used to have a horribly written node process that was running in a Mesos cluster (using Marathon). It had a memory leak and would start to fill up memory after about a week of running, depending on what customers were doing and if they were hitting it enough. The solution, rather than investing time in fixing the memory leak, was to add a cron job that would kill/reset the process every three days. This was easier…

I once managed a cluster of worker servers with an @reboot cronjob that scheduled another reboot after $(random /8..24/) hours. They took jobs from a rabbitmq queue and launched docker containers to run them, but had some kind of odd resource leak that would lead to the machines becoming unresponsive after a few days. The whole thing was cursed honestly but that random reboot script got us through for a few more years until it could be replaced with a more modern design.

Re: Why Twitter didn’t go down: From a real Twitter SRE

#920
post #648
post #515

Earlier quoted context omitted.

> "Anyone who actually writes software, please report to the 10th floor at 2 pm today. Before doing so, please email a bullet point summary of what your code commands have achieved in the past ~6 months, along with up to 10 screenshots of the most salient lines of code" Actual quote. Anyone using the term "code commands" comes out a little detached from programming reality, let alone the rest of this request, it is o…

"Code commands" is very plausibly an autocomplete flub of what was supposed to be "code commits." When I type "code comm" my iPhone offers up "commands" as the completion. I've seen a lot of mockery of this request, but I suspect people aren't considering the wide variance in employee quality that can exist within a mismanaged organization. What Musk was asking for here wouldn't be a good way to evaluate skilled, con…

What about "most salient lines of code"?
Post reply on HN