Live data from Hacker News

Why Twitter didn’t go down: From a real Twitter SRE

matthewtejo.substack.com

391–400 of 1001 posts

Re: Why Twitter didn’t go down: From a real Twitter SRE

#391

I think the real question is: Twitter grew 3x on the headcount front with a flat stock price over the course of less than 5 years. What exactly where these thousands of employees actually doing and why did the previous CEO think what they were doing was worth hiring them for? That's just basic accountability from a stock holder or employee perspective. That's apparently a ton of money being wasted on nothing at all.

Twitter used to experience significant downtime compared to all other major platforms and one of the reason was its lack of redundancies across everything. Headcount is one such thing and it takes manpower to automate infrastructures as discussed in the post. Sure, you can run the platform with 1/10 headcount with significantly degraded user experiences (say ~98%). This is not a problem for startups but people usuall…

I think the opposite. Many softwares at its best when the team was small. Software companies have to hire many people because it needs to report growth to investors, headcount is one of the measurement of growth. It is not necessarily good for the product, actually many times, it hurts the product, but overall it is good for the company, the company will enter new areas, can explore new things.

What Twitter is doing is to scale down first, focus on the product, and once it gains traction, it definitely can scale up again. I don't think it will hurt the product very much.

Re: Why Twitter didn’t go down: From a real Twitter SRE

#392

You can usually do more with less if the remaining few are up to the challenge. It will come at some cost, like having to "move fast and break things" more. There can even be a positive side if the previous situation had a lot of friction due to "too many cooks," as I've experienced on several teams. What's more worrisome is the alleged 80hr work weeks. Yes, two SWEs on a team aren't as good as one SWE working double…

> cut corners on actually verifying the subscribers

Did they actually verify anything other than the ability to pay $8? It seems wild to me that they thought it would work out just fine

Re: Why Twitter didn’t go down: From a real Twitter SRE

#393
post #158

Tangent but I do hope that Musks trim down results in orgs that have less “executives” and a layered cake of a org structure, and more autonomous small teams that execute on shared overarching initiatives. I really don’t understand why so many tech companies have like 8 layers of engineering levels. If the argument is that you need more money so more levels, just have a bigger band. Don’t chase titles they don’t mean…

> Tangent but I do hope that Musks trim down results in orgs that have less “executives” and a layered cake of a org structure, and more autonomous small teams that execute on shared overarching initiatives.

The core problem is that this approach doesn't really scale out because communication overhead exhibits quadratic growth to the orgs size if it's untamed. The feasible options are:

  1. Let the complexity bring chaos across the org
  2. One decision maker rule them all
  3. Gives some sort of management structures to the org
Your proposal is somehow between the option 1 and 2. The option 1 works pretty well for smaller orgs and it might scale to a quite sizable business if the members are generally competent so org-wide trust can be well-established. But anyway you'll hit a road blocker eventually since people cannot spend all of their time on communication overheads. The option 2 just moves the burden of entire complexity into a single personnel so it's not really a reproducible solution but more of a mere luck.

Hence the option 3 is the only remaining option for regular orgs and many smart people tried to figure out the best structure (or at least best practices) but unfortunately we don't have a definitive answer yet. Google-style "tech level" is one of the tool to reduce communication overhead by setting a common structure for expectations (e.g. "we have 1 L6 and 3 L5s to take that project" is generally easier to convince than length explanations of your team members). It's not ideal but it somehow works so it's adopted.

You're likely right that you'll be much more productive if you can get rid of those bureaucracies, but getting other folks convinced is a completely different story. Trust takes time to propagate and people have a limited time to spend on it. This obviously could be drastically simplified if you can work with Elon (or similar style leaders) directly but his time is extremely limited so there will always be only a small number of people who can enjoy that privilege...

Re: Why Twitter didn’t go down: From a real Twitter SRE

#394
post #284

Earlier quoted context omitted.

It's actually scary how many people, even engineers, put their reputation on the line saying Twitter wouldn't survive the weekend. It wasn't just Twitter employees. It's like a mass psychosis of some kind. It comes off as a kind of desperation, as though they need Elon to fail. Why? What's driving that response?

There are a lot of places where the systems would start to fall over if 70-80% of the team departed. Especially since a lot of folks left on bad terms and/or were suddenly terminated. It was the opposite of a smooth handover. So it wasn't unreasonable to think that Twitter would begin experiencing problems. I didn't think that Twitter was going to literally have an unrecoverable system crash and permanently shut its…

Just not like him yes, but also probably feel some desire to defend their profession.

Re: Why Twitter didn’t go down: From a real Twitter SRE

#395

The reason why Twitter hasn’t crashed is it’s well written and well orchestrated. Once a bug comes in and crashes something that is when the chaos starts. It’s almost a guarantee that the current crop do not know how to fix the bug. It will be interesting to see how they handle that.

> It’s almost a guarantee that the current crop do not know how to fix the bug

Seeing lots of comments like this. Why do you feel it’s ok to say that, when you don’t actually work there or have intimate knowledge of who is currently working there? The pure speculation and nonsense surrounding Twitter at the moment is plain awful.

Re: Why Twitter didn’t go down: From a real Twitter SRE

#396

Now that Elon has cleaned house, when is Twitter hiring again? It might be a good opportunity to get in on the new ground floor.

Why would anyone want to work in a dumpster fire?

I've wondered if the hiring funnel will be so lean that the TC may be forced high

Re: Why Twitter didn’t go down: From a real Twitter SRE

#397
post #158

Tangent but I do hope that Musks trim down results in orgs that have less “executives” and a layered cake of a org structure, and more autonomous small teams that execute on shared overarching initiatives. I really don’t understand why so many tech companies have like 8 layers of engineering levels. If the argument is that you need more money so more levels, just have a bigger band. Don’t chase titles they don’t mean…

> If the argument is that you need more money so more levels, just have a bigger band.

Someone earning a lot more than you at the same nominal position leads to a lot of resentment: the perception is that you are clearly wronged here. On the other hand a rank system makes this less objectionable and offers at least some roadmap to a similar income. "Oh, she's SWE L9000 naturally she'd earn that"

Re: Why Twitter didn’t go down: From a real Twitter SRE

#398

Now that Elon has cleaned house, when is Twitter hiring again? It might be a good opportunity to get in on the new ground floor.

How would stock grants work, cause I doubt he'll IPO twitter? Or would it just be netflix-style cash offers

Re: Why Twitter didn’t go down: From a real Twitter SRE

#399

I think the real question is: Twitter grew 3x on the headcount front with a flat stock price over the course of less than 5 years. What exactly where these thousands of employees actually doing and why did the previous CEO think what they were doing was worth hiring them for? That's just basic accountability from a stock holder or employee perspective. That's apparently a ton of money being wasted on nothing at all.

> Twitter grew 3x on the headcount front There were multiple executives making $10m/yr+ There were board members There were shareholders Why did all of them not stop this headcount increase if it's as easily reduced as "too much headcount bad, smaller headcount good"? These are paid professionals who are supposedly wealthy, good at their jobs, smart, informed, etc. How can us commenters on HackerNews sit from our arm…

I've worked in multiple financial services companies where management is incentivised to be as ruthless as possible and they are always overstaffed in areas and understaffed in others. I've been in teams of 10 people that could be staffed by 2.

Hiring often isn't done because of current requirements. Senior execs come and go and with them so do strategic objectives. You accumulate people and they're often not laid off when the thing they work on becomes redundant. Large scale layoffs are awful for morale and usually only come after a 'crisis' occurs.

Re: Why Twitter didn’t go down: From a real Twitter SRE

#400

From this operations engineer's perspective, there are only 3 main things that bring a site down: new code, disk space, and 'outages'. If you don't push new code, your apps will be pretty stable. If you don't run out of disk space, your apps will keep running. And if your network/power/etc doesn't mysteriously disappear, your apps will keep running. And running, and running, and running. The biggest thing that brings…

> absolutely bring a site down over time, is expired certs From today's Casey Newton's newsletter: In early December, a number of Twitter’s security certificates are set to expire — particularly those that power various back-end functions of the site. (“Certs,” as they are usually called, serve to reassure users that the website they are visiting is authentic. Without proper certs, a modern web browser will refuse to…

I can imagine both cases being true, that the renewal process is automated and that certs won't get renewed because institutional knowledge has left the door. Where I'm at, service-to-service TLS certificates (the bulk of our certs) are automatically rotated by our deploy systems. But there are always the edge cases: the certificates manually created a long time ago (predating any standardized monitoring systems) with long expiry dates, and certificates for systems that simply can't run off the standard infrastructure. Sometimes, they'll bring down systems with low SLOs; other times, they'll block all internal development.
Post reply on HN