Live data from Hacker News

Why Twitter didn’t go down: From a real Twitter SRE

matthewtejo.substack.com

871–880 of 1001 posts

Re: Why Twitter didn’t go down: From a real Twitter SRE

#872
post #800
post #81

Earlier quoted context omitted.

A few parts clearly did go down. 2FA login was just serving error codes all day a few days ago. On Saturday people were posting entire feature films, 2 minutes per tweet, because apparently their copyright content matching wasn’t working.

I would use it as a case in point. You have one group claiming that Twitter is completely on fire and soon to be closed, and the other saying it's never been better and there was zero fallout from rapidly firing thousands of people including a lot of engineers. Neither of those extreme viewpoints reflected reality accurately. There were significant problems, but they did not take the entire site down or anything. In…

I think both sides are correct. Anyone with experience at the scale knows that if everyone stops working, the site becomes more reliable. Holiday freezes demonstrate this. But holiday freezes also demonstrate the other thing: if left alone for too long the systems start to rot, from memory leaks and cache ossification and other things that usually aren't noticed during active development.

Personally, I doubt that it will be anything technological that ends Twitter. It will be economic. Their advertising revenue has been decimated and their operating costs have never been higher.

Re: Why Twitter didn’t go down: From a real Twitter SRE

#873

Earlier quoted context omitted.

> removing engineers won't instantly crash the product. It'll happen slowly It's amazing to me how many people following the Twitter saga, some familiar with or actually working in technology, thought that Twitter would crash within days of the engineers being fired. And because it didn't, the job cuts are justified.

Quoted post unavailable.

Weird to hear the work culture that involves being nice to people being called the toxic one.

Re: Why Twitter didn’t go down: From a real Twitter SRE

#874

Earlier quoted context omitted.

> In a start-up many of these roles can be ignored, becaus growth > stability. In a large organization, part of the bloat helps insure a certain amount of stability that's necessary to keep an organization alive. It also (a) increases the bus factor, [1] and (b) allows people to take vacations and time off without having to watch their phones like hawk. [1] https://en.wikipedia.org/wiki/Bus_factor

Good point. I know this as the Mack Truck Theory. For the project I'm working on right now, there's a couple of incredibly valuable people that would cause a pretty significant issue if they disappeared.

Sensitive workplaces (which seems like most these days) have taken to calling this the Lottery Factor (as in team members quit bc they hit the lottery) to spare more delicate types the pain of imagining their peers run over in traffic accidents.

Re: Why Twitter didn’t go down: From a real Twitter SRE

#875

Earlier quoted context omitted.

I agree those were odd takes. I've likened firing most of the engineers to taking your hands off the wheel in the car. It won't crash immediately, but it doesn't mean the car can go driverless. With that said, there are differences between internal systems and something like Twitter on the public internet. I assume that Twitter is a system under constant attack. What happens when the next log4shell level vulnerabilit…

If Twitter went another month without an outage, how would you adjust your opinion? How about a year? The car analogy is amusing, but how much does it really hold up? Have we ever seen another major social media company drop this much of its staff in one go? I certainly can’t think of an example. I think we’re in somewhat uncharted waters here. A driverless car won’t last long, we know that for a fact. I think it rem…

> I think it remains to be seen how long a bloatless twitter can last.

I'm not convinced Twitter had a ton of bloat. (Most of the teams actually involved don't seem to think so). Just because Elon can't understand something, doesn't make that thing "bloat".

Twitter definitely had a few weird features that could be cut (the audio podcasting thing, for example). But calling most of Twitter microservices "bloat" is about as dumb as calling a cars Seatbelt and Airbag and Crumple Zones and that spare tire in the trunk "bloat" -- it's only "bloat" if you assume all people will always be perfect and no one will ever make a mistake anywhere, and nothing bad will ever happen.

Re: Why Twitter didn’t go down: From a real Twitter SRE

#876

> This left a lot wondering what exactly was going on with all those engineers and made it seem like it was all just bloat. I was partly expecting the rest of the article to explain to me why exactly it wasn't just bloat. But it goes on talking about this 1~3-person cache SRE team that built solid infra automation that's really resilient to both hardware and software failures. If anything, the article might actually…

> the article might actually persuade me that it was all bloat First of all, how does it persuade you of that? The article touches a really small (though incredibly important for up-time) subject. Secondly, in any large company, the majority is 'bloat'. It's security engineers, code reviews, data architecture, HR, internal audit teams, content moderators, ccrum masters and I can keep going. In a start-up many of thes…

The whole point of the article is twitter was designed to be resilient. (and it shows, twitter has great uptime). And the whole point of resiliency, beyond not negatively impacting customer experience, is to buy engineers time to fix things when stuff breaks.

What we are watching is a massive failure event right now and the question really is if there's enough time for twitter management to fill in the gaps before there's an outage.

Re: Why Twitter didn’t go down: From a real Twitter SRE

#877

Earlier quoted context omitted.

Good point. I know this as the Mack Truck Theory. For the project I'm working on right now, there's a couple of incredibly valuable people that would cause a pretty significant issue if they disappeared.

Sounds like those people are in strong negotiation positions

H-1B would seem to belie that.

Re: Why Twitter didn’t go down: From a real Twitter SRE

#878

Earlier quoted context omitted.

> removing engineers won't instantly crash the product. It'll happen slowly It's amazing to me how many people following the Twitter saga, some familiar with or actually working in technology, thought that Twitter would crash within days of the engineers being fired. And because it didn't, the job cuts are justified.

I'm still on the fence. As an engineering manager, I tend to attach faces to those "jobs". So seeing cuts, I imagine a ton of people that had to go home and tell their friends/family/etc that they no longer had a job. On the other hand. As an engineer, we tend to attach way too much self importance to our roles. Like if we're not there entering the "numbers" 4 6 15 16 24 32 every 108 minutes, the entire business is g…

You got the numbers wrong and Elon fired the person who knows them...

Re: Why Twitter didn’t go down: From a real Twitter SRE

#879

Earlier quoted context omitted.

> If the current state of affairs at Twitter keeps up, it'll probably be a slow descent into chaos. It's not like Twitter was bug free before. How many times it annoyingly refreshed the timeline while I was reading something, or when it shows notification that it failed to send the DM, and when you retry it says "you've already wrote this", or you open the reply dialog, but it freezes, has no send button at all, so y…

They basically gave an established global corporation with around 200mn active users a startup tech, marketing, and support stack. Anyone thinking "huge success" is unrealistically optimistic, IMO. It's going to be another MySpace/AOL/Bebo, added to the list of dumbest purchases ever. And that's still going to be true if the point was to destroy the original community and replace it with a different political orienta…

You're right on this potentially being the worse purchase of all time. The purchase price of 44 Billion was overvalued. If the only thing Twitter is supposed to do is to move tweets around then there was a ton of bloat of developers with a massively bloated purchase price. If Twitter had quality R&D products that could compete with TikTok, etc., then there may not have been a bloated workforce nor bloated purchase price.

Re: Why Twitter didn’t go down: From a real Twitter SRE

#880

Earlier quoted context omitted.

> removing engineers won't instantly crash the product. It'll happen slowly It's amazing to me how many people following the Twitter saga, some familiar with or actually working in technology, thought that Twitter would crash within days of the engineers being fired. And because it didn't, the job cuts are justified.

I'm still on the fence. As an engineering manager, I tend to attach faces to those "jobs". So seeing cuts, I imagine a ton of people that had to go home and tell their friends/family/etc that they no longer had a job. On the other hand. As an engineer, we tend to attach way too much self importance to our roles. Like if we're not there entering the "numbers" 4 6 15 16 24 32 every 108 minutes, the entire business is g…

It's great to hear your concern with the actual people. That tends to get lost for some reason. And I agree we tend to attach too much importance to our roles, but the flip side bears a certain amount of truth, though on a timeline more like 108 days than 108 minutes.
Post reply on HN