Live data from Hacker News

Why Twitter didn’t go down: From a real Twitter SRE

matthewtejo.substack.com

881–890 of 1001 posts

Re: Why Twitter didn’t go down: From a real Twitter SRE

#881

Earlier quoted context omitted.

The headcount at WhatsApp in 2013 was somewhere between 50-100, at which time they were servicing approx 400m MAU, which is more than users than Twitter has been able to boast for most of their existence. Coincidentally, in 2013 SpaceX was just starting to provide commerical launch capacity, at which point I think they too had Not surprised Elon Musk thinks he can run twitter with a skeleton crew.

1) And what was their uptime in 2013? How did uptime change as the service grew in popularity? 2) WhatsApp does not support the type of public broadcasts done at Twitter, and due to its e2ee doesn't require much human moderation.

1. WhatsApp was more reliable before migration to Meta's infrastructure.

2. WhatsApp didn't have e2ee back then. The broadcasting is important, yes, but it is very heavily biased towards reads over writes, so something like Cloudflare would solve 99% of the load.

Re: Why Twitter didn’t go down: From a real Twitter SRE

#882

Earlier quoted context omitted.

> removing engineers won't instantly crash the product. It'll happen slowly It's amazing to me how many people following the Twitter saga, some familiar with or actually working in technology, thought that Twitter would crash within days of the engineers being fired. And because it didn't, the job cuts are justified.

I'm still on the fence. As an engineering manager, I tend to attach faces to those "jobs". So seeing cuts, I imagine a ton of people that had to go home and tell their friends/family/etc that they no longer had a job. On the other hand. As an engineer, we tend to attach way too much self importance to our roles. Like if we're not there entering the "numbers" 4 6 15 16 24 32 every 108 minutes, the entire business is g…

> Like if we're not there entering the "numbers" 4 6 15 16 24 32 every 108 minutes, the entire business is going to crumbl

Never I have encountered an engineer that thought that.

Re: Why Twitter didn’t go down: From a real Twitter SRE

#883
post #328
post #298

Earlier quoted context omitted.

One SRE, many SWE. Also have fun asking someone to be permanently oncall with one person on the team. The cache clusters size are also described here for anyone who wants a good technical read over speculation. https://www.usenix.org/system/files/osdi20-yang.pdf

The OP claims he did the implementation (so he was the software engineer too?): > I designed and implemented most of the tools that are keeping it running so I think I’m qualified to talk about it.

I read this as they built the “tools” (automation, orchestration, monitoring, etc.) for this system, not the system itself; which aligns with the common definition of SRE.

Re: Why Twitter didn’t go down: From a real Twitter SRE

#884

Earlier quoted context omitted.

Except in these firing sprees it's not competent employees that remain - those people are usually the first to jump ship because they have options and see where it's headed. It's not like they get raises for staying to get a part of the savings from the cuts - usually they get pay freezes. Restructuring is done on whatever idea the new owner(s) have in plan - which could be equally disconnected from "real product". I…

You realise how annoying it is when you see people around you in the company making as much money as you and doing nothing useful to actually generate revenue. Data scientists that write a blog post about the most used hashtag or emoticon. Or PMs that we don't need... I'm staying and I'm glad the teams are trimmed.

Are you absolutely 100% sure the surface stuff you had visibility into was all they did?

Re: Why Twitter didn’t go down: From a real Twitter SRE

#885

Earlier quoted context omitted.

Quoted post unavailable.

> The job cuts are clearly justified because of the extremely toxic "work culture" I'm very curious in what way the stuff Elon is doing is not creating toxic work culture? >Minimum I'm not denying there may have been toxic work culture before, but I don't see how replacing it with another form of toxic work culture is to be celebrated. Also, writing this: > There is no active independent thought around this subject t…

[deleted]

Re: Why Twitter didn’t go down: From a real Twitter SRE

#886

Earlier quoted context omitted.

Quoted post unavailable.

> The job cuts are clearly justified because of the extremely toxic "work culture" I'm very curious in what way the stuff Elon is doing is not creating toxic work culture? >Minimum I'm not denying there may have been toxic work culture before, but I don't see how replacing it with another form of toxic work culture is to be celebrated. Also, writing this: > There is no active independent thought around this subject t…

I am not trying to comment on anything else here, but this:

>>MinimumThose are the norm for most fulltime people in most jobs. You know that, right? Right?

Re: Why Twitter didn’t go down: From a real Twitter SRE

#887
post #326
post #27

All this does is point out that smart people worked at Twitter who may now no longer work there, whether on their own accord, or due to Elon’s bulldogging tactics. Elon thinks he knows what he’s doing, but what he is going to be left with are people who are willing to work hard by his standards, but not necessarily smart. The simple truth is Elon knows nothing about the actual work involved in tech. He knows words or…

John Carmack, "Elon is definitely an engineer. He is deeply involved with technical decisions at spacex and Tesla. He doesn’t write code or do CAD today, but he is perfectly capable of doing so." Kevin Watson, who developed the avionics for Falcon 9 and Dragon and previously managed the Advanced Computer Systems and Technologies Group within the Autonomous Systems Division at NASA's Jet Propulsion laboratory: "Elon i…

Adding another datapoint from one of my previous comments:

In response to someone saying on Twitter how Elon doesn't understand the technical stuff of rocketry, Tom Meuller, former CTO of Propulsion at SpaceX and the designer of many of their engines responded

"I worked for Elon directly for 18 1/2 years, and I can assure you, you are wrong"

https://twitter.com/lrocket/status/1512919230689148929?s=20&...

Re: Why Twitter didn’t go down: From a real Twitter SRE

#888

Earlier quoted context omitted.

> removing engineers won't instantly crash the product. It'll happen slowly It's amazing to me how many people following the Twitter saga, some familiar with or actually working in technology, thought that Twitter would crash within days of the engineers being fired. And because it didn't, the job cuts are justified.

I'm still on the fence. As an engineering manager, I tend to attach faces to those "jobs". So seeing cuts, I imagine a ton of people that had to go home and tell their friends/family/etc that they no longer had a job. On the other hand. As an engineer, we tend to attach way too much self importance to our roles. Like if we're not there entering the "numbers" 4 6 15 16 24 32 every 108 minutes, the entire business is g…

If you think the people you work with don't have a sincere desire to accomplish things you are a terrible manager.

Re: Why Twitter didn’t go down: From a real Twitter SRE

#889

> This left a lot wondering what exactly was going on with all those engineers and made it seem like it was all just bloat. I was partly expecting the rest of the article to explain to me why exactly it wasn't just bloat. But it goes on talking about this 1~3-person cache SRE team that built solid infra automation that's really resilient to both hardware and software failures. If anything, the article might actually…

Argh. "It works now, so it will work until forever." It takes _effort_ to make it work this smoothly now, _and in the future_. SRE is about _preventing_ issues. Not mopping up after them. To me, the article read like every succesfull sysadmin story: there's no fires, so sysadmin must be bloat.

To be honest I was very surprised to hear what a cache SRE was working on. It sounded like he had to build all of handling of hardware issues, rack awareness and other basic datacenter stuff himself. Does it mean that every specialized team also had to do it? Why would cache engineer need to know about hardware failures at all, its datacenter team's responsibility to detect and predict issues and shutdown servers gracefully if possible. It should be completely abstracted from cache SRE, like cloud abstracts you from it. Yet he and is team spends years on automation around this stuff using Mesos stack that they probably regret adopting by now. I feel like in this zoomed in case of twitter caches what they were working on is questionable, but the team size seems to be adequate to the task, so my takeaway is that like any older, larger company Twitter accumulated fair amount of tech debt and there is no one to take large scale initiative to eliminate it.

Re: Why Twitter didn’t go down: From a real Twitter SRE

#890
post #162

From this operations engineer's perspective, there are only 3 main things that bring a site down: new code, disk space, and 'outages'. If you don't push new code, your apps will be pretty stable. If you don't run out of disk space, your apps will keep running. And if your network/power/etc doesn't mysteriously disappear, your apps will keep running. And running, and running, and running. The biggest thing that brings…

Another thing we noticed at Netflix was that after services didn’t get pushed for a while (weeks), performance started degrading because of things like undiscovered memory leaks, threads leaks, disks filling up. You wouldn’t notice during normal operations because of regular autoscaling and code pushes, but code freezes tended to reveal these issues.

We joked about adding this to the NodeQuark platform:

    // Fix Slow Memory Leaks
    setTimeout(() => process.exit(1), 1000 * 60 * 60 * 24)
Post reply on HN