Live data from Hacker News

Why Twitter didn’t go down: From a real Twitter SRE

matthewtejo.substack.com

581–590 of 1001 posts

Re: Why Twitter didn’t go down: From a real Twitter SRE

#581

> This left a lot wondering what exactly was going on with all those engineers and made it seem like it was all just bloat. I was partly expecting the rest of the article to explain to me why exactly it wasn't just bloat. But it goes on talking about this 1~3-person cache SRE team that built solid infra automation that's really resilient to both hardware and software failures. If anything, the article might actually…

> the article might actually persuade me that it was all bloat First of all, how does it persuade you of that? The article touches a really small (though incredibly important for up-time) subject. Secondly, in any large company, the majority is 'bloat'. It's security engineers, code reviews, data architecture, HR, internal audit teams, content moderators, ccrum masters and I can keep going. In a start-up many of thes…

> If the current state of affairs at Twitter keeps up, it'll probably be a slow descent into chaos.

It's not like Twitter was bug free before. How many times it annoyingly refreshed the timeline while I was reading something, or when it shows notification that it failed to send the DM, and when you retry it says "you've already wrote this", or you open the reply dialog, but it freezes, has no send button at all, so you have to re-open it. All of this was happening to me pretty regularly long before Elon came along.

As we all know, just hiring more people is not necessarily the solution to every problem, and to me it seems it was exactly what Twitter tried to do in the past. Now they deconstructed it to the bare bones, which will clearly show what are the core problems and requirements. They basically turned Twitter back into a startup. And from that new starting point they can hire again to cover the needs as they arise. If they succeed it will be a huge success as they'll end up with far more optimal team (and huge savings), and of course, if they fail to catch up with problems it will be a huge failure. We'll see how well Musk can manage it...

Re: Why Twitter didn’t go down: From a real Twitter SRE

#582

I don't think it will fail on a technical level. As this article says lots of engineering has gone in to make the thing pretty resilient. I would also say that there are still enough engineers who work there who can figure out what's gone wrong and "turn it on and off again" or w/e makes it splutter back to life. In terms of changes to the Platform ditto. It's not difficult to make these changes that a team of 100's…

Are you saying Twitter is paying ~$1.5B in interest rates every year?

I hear $1bn quoted but maybe $1.5nm is accurate.

Re: Why Twitter didn’t go down: From a real Twitter SRE

#583
post #27

All this does is point out that smart people worked at Twitter who may now no longer work there, whether on their own accord, or due to Elon’s bulldogging tactics. Elon thinks he knows what he’s doing, but what he is going to be left with are people who are willing to work hard by his standards, but not necessarily smart. The simple truth is Elon knows nothing about the actual work involved in tech. He knows words or…

You know, I think Musk is an ass, and would never work for him, but don't you think that someone who has managed to launch and then run many successful and complex technology projects might actually know a thing or two about launching and running simpler technology projects? And if you're going to claim that his successes have been due to the people surrounding him who actually know what they are doing, then all that…

Well, I think the issue is precisely considering Twitter a "simple technology project", and it's the same mistake that Musk does. Twitter isn't a "software and servers business" as he said. Twitter is a social community, and while in some regards it might be easier, it's also far more difficult in others. Just compare how many business and institutions can reliably launch rockets or create cars, and how many can reliably create social networks.

Re: Why Twitter didn’t go down: From a real Twitter SRE

#584

Earlier quoted context omitted.

Automotive and aerospace are not that similar to social media. People buying into the vision of "get the planet off fossil fuels for transport" and "get this species to Mars" are probably willing to make sacrifices that people working on social media are not. It's the Halo Effect fallacy to think competence in one field automatically translates to another. Especially when the founder in question has displayed increas…

Automotive and aerospace are not similar to each other, either. I don't know any other outfit that was successful at both. > It's the Halo Effect fallacy to think competence in one field automatically translates to another I didn't say it was. I was responding the notion that Musk blundered into success at Tesla and SpaceX.

> Automotive and aerospace are not similar to each other, either. I don't know any other outfit that was successful at both.

Rolls-Royce (of the last century) would qualify, but it was more aero than space.

Re: Why Twitter didn’t go down: From a real Twitter SRE

#585

Earlier quoted context omitted.

> I was partly expecting the rest of the article to explain to my why exactly it wasn't just bloat Same here. I guess his header was on point in why Twitter is still up; but I was also interested in hearing about why Twitter actually needs all those people. If it can be run with 50-80% of the staff gone, that does sound like some bloat at least.

Slack space leads to innovations, like developing infrastructure automation and improving capacity planning. SRE as a practice needs slack space for operations teams to work on improvements and fixes in addition to BAU fault fixing, deployments and patching.

> Slack space leads to innovations

I'm certainly no supporter of "lean operations" with minimum staff etc, and fully agree that you need people that are well rested (for the lack of a better analogy) to do great stuff. But I do think that some of these internet giants do have to many people working there; wasn't LinkedIn 14_000 strong when Microsoft bought it?

I've always felt that the American model of doing business is based on how we optimize network traffic, i.e. double the amount of data until failure; then turn in down a bit. Fire on all cylinders until people are truly worn out, then turn down the pace a bit. Haven't worked in the US so I'm pulling this info out of thin air...

Re: Why Twitter didn’t go down: From a real Twitter SRE

#586
post #320

I think the real question is: Twitter grew 3x on the headcount front with a flat stock price over the course of less than 5 years. What exactly where these thousands of employees actually doing and why did the previous CEO think what they were doing was worth hiring them for? That's just basic accountability from a stock holder or employee perspective. That's apparently a ton of money being wasted on nothing at all.

Did't you see the "Day in my life at the Twitter office video!"? https://www.tiktok.com/@realpankhilpatel/video/7159187292631... Normal people don't have vacations like that.

There is nothing here that other big tech companies don't have. To attract the best, they spend a shitload on perks and benefits.

Re: Why Twitter didn’t go down: From a real Twitter SRE

#587

Earlier quoted context omitted.

> the article might actually persuade me that it was all bloat First of all, how does it persuade you of that? The article touches a really small (though incredibly important for up-time) subject. Secondly, in any large company, the majority is 'bloat'. It's security engineers, code reviews, data architecture, HR, internal audit teams, content moderators, ccrum masters and I can keep going. In a start-up many of thes…

data architecture is bloat? Have you implemented a system which stores hundreds of billions of pieces of media content and makes different slices of them immediately available to hundreds of millions of users?

No but the people who made clickhouse did. I don’t think 1000s of engineers were required for that.

Where’s the 1000s of engineers for Postgres? Most stuff that works is made by a handful of people. Look at io_uring it’s basically one guy at Facebook…

Re: Why Twitter didn’t go down: From a real Twitter SRE

#589
post #451

From this operations engineer's perspective, there are only 3 main things that bring a site down: new code, disk space, and 'outages'. If you don't push new code, your apps will be pretty stable. If you don't run out of disk space, your apps will keep running. And if your network/power/etc doesn't mysteriously disappear, your apps will keep running. And running, and running, and running. The biggest thing that brings…

It's crazy to think about, but many people who use and build software today, including HN readers/commenters, are young enough to have only been exposed to the SaaS, cloud-first era, where software built with microservices deployed from CI/CD systems multiple times per day is just the way things are done. You're totally right; if you don't make changes to the software, it's unlikely to spontaneously stop working, esp…

The other side of that is browsers. Even if you don’t change your code, the platform people are running your code in changes, automatically in many cases. New JS or CSS behavior in next safari or chrome? You need to patch/push to accommodate running environments that are outside your control.

Re: Why Twitter didn’t go down: From a real Twitter SRE

#590

I don't think it will fail on a technical level. As this article says lots of engineering has gone in to make the thing pretty resilient. I would also say that there are still enough engineers who work there who can figure out what's gone wrong and "turn it on and off again" or w/e makes it splutter back to life. In terms of changes to the Platform ditto. It's not difficult to make these changes that a team of 100's…

Are you saying Twitter is paying ~$1.5B in interest rates every year?

It’s going to be something like that. He bought it for 44 billion. That was mostly loans. Those loans are transferred to the company and the company pays interest on them.
Post reply on HN