Live data from Hacker News

Why Twitter didn’t go down: From a real Twitter SRE

matthewtejo.substack.com

361–370 of 1001 posts

Re: Why Twitter didn’t go down: From a real Twitter SRE

#361

From this operations engineer's perspective, there are only 3 main things that bring a site down: new code, disk space, and 'outages'. If you don't push new code, your apps will be pretty stable. If you don't run out of disk space, your apps will keep running. And if your network/power/etc doesn't mysteriously disappear, your apps will keep running. And running, and running, and running. The biggest thing that brings…

> you don't run out of disk space (from logs for example)

For a social media / user-generated content application, the macro storage concerns are a lot more important than the micro ones. By this I mean, care more about overall fleet-wide capacity for product DBs and media storage, instead of caring about a single server filling up its disk with logs.

With UGC applications, product data just grows and grows, forever, never shrinking. Even if the app becomes less popular over time, the data set will still keep growing -- just more slowly than before.

Even if your database infrastructure has fully automated sharding, with bare metal hosting you still need to keep doing capacity planning and acquiring new database hardware. If no one is doing this, it's game over, there's simply nowhere to store new tweets (or new photos, or whichever infra tier runs out of hardware first...)

Staffing problems in other eng areas can exacerbate this. For example, if automated bot detection becomes inadequate, bot posting volume goes way up and takes up an increasing amount of storage space.

Re: Why Twitter didn’t go down: From a real Twitter SRE

#362

I think people's expectations became so exaggerated it was inevitable they wouldn't be lived up to. I'm sure Twitter will experience degradation from the drastic cost-cutting, but it was never going to happen overnight and I'm not sure why news outlets were saying that (except that their sources were employees with slightly inflated senses of their own importance, which we're all sometimes guilty of). And people beca…

I did see a lot of politically-motivated salivating which I thought was woefully premature. That said, I see 2 outcomes now: 1. Generally, with large, complex systems like this, everything works, until it doesn't. All the big boys have major outages periodically. I just can't fathom how Twitter is going to handle the eventual certainty of a major outage when, as the author notes, in some cases there are teams that ha…

I also want to add a #3: the crew that's left is probably on-call 24/7. My thoughts are with the poor souls on that rotation (if your team even still has a "rotation")

Re: Why Twitter didn’t go down: From a real Twitter SRE

#363

Earlier quoted context omitted.

Automotive and aerospace are not that similar to social media. People buying into the vision of "get the planet off fossil fuels for transport" and "get this species to Mars" are probably willing to make sacrifices that people working on social media are not. It's the Halo Effect fallacy to think competence in one field automatically translates to another. Especially when the founder in question has displayed increas…

Automotive and aerospace are not similar to each other, either. I don't know any other outfit that was successful at both. > It's the Halo Effect fallacy to think competence in one field automatically translates to another I didn't say it was. I was responding the notion that Musk blundered into success at Tesla and SpaceX.

What I think makes people skeptical of him is that deep down we all believe in nominative determinism. Who do you think runs a rocket company behind the scenes, "Shotwell" or "Musk"?

Re: Why Twitter didn’t go down: From a real Twitter SRE

#364

Earlier quoted context omitted.

Nice work! Does Twitter runs their own DCs or hosted somewhere else?

As of three years ago, Twitter had 2 to 3 data centers, and was moving some stuff to GCP. Not sure if current state of things.

>Mr. Musk is also considering shuttering one of Twitter’s three main U.S. data centers, a location known as SMF1 in Sacramento, which is used to store information needed to run the social media site, four people with knowledge of the effort said. If the data center in Sacramento is taken offline, it will leave the company with data centers in Atlanta and Portland, Ore., with potentially less backup computing capacity in case something fails.

https://www.nytimes.com/2022/11/18/technology/elon-musk-twit...

Re: Why Twitter didn’t go down: From a real Twitter SRE

#365
post #81
post #43

It was if I was in two realities the other day. I did experience a few bugs but twitter never went "down" for me, despite everyone claiming that 1) it would and 2) it already did. All 10 of the "top trends" for me were about how twitter was dead/dying and armchair experts on reddit said it was only a matter of hours at this point. Maybe it was close (in reference to external factors not implied in your very insightfu…

A few parts clearly did go down. 2FA login was just serving error codes all day a few days ago. On Saturday people were posting entire feature films, 2 minutes per tweet, because apparently their copyright content matching wasn’t working.

It's hard to really say whether that was due to something Elon did or if it was just a normal outage from code deployments, hardware failures, etc. Anything that might seem like an outage will be magnified now because everyone will think it was Elon's doing and it will get tons of clicks. The reality is that these huge sites have a dozen small/medium outages almost every day, but almost no one notices.

Re: Why Twitter didn’t go down: From a real Twitter SRE

#366
post #162

From this operations engineer's perspective, there are only 3 main things that bring a site down: new code, disk space, and 'outages'. If you don't push new code, your apps will be pretty stable. If you don't run out of disk space, your apps will keep running. And if your network/power/etc doesn't mysteriously disappear, your apps will keep running. And running, and running, and running. The biggest thing that brings…

Another thing we noticed at Netflix was that after services didn’t get pushed for a while (weeks), performance started degrading because of things like undiscovered memory leaks, threads leaks, disks filling up. You wouldn’t notice during normal operations because of regular autoscaling and code pushes, but code freezes tended to reveal these issues.

Agreed, one of the craziest bugs I had to deal with was we had a distributed system using lots of infrastructure. Said distributed system started having trouble communicating with random nodes and sub-systems. I spent 3 hard days finding a Linux kernel bug where the ARP cache was not removing least recently accessed network addresses. Normally, this wouldn't be a big deal for a typical network because few networks would fill up the default arp cache size. That was even true for ours except that we would slowly add and remove infrastructure over the course of a couple months until eventually the ARP cache would fill and remove the random network devices... It wasn't even our distributed application code... Some bugs take time to manifest themselves in very creative ways.

Re: Why Twitter didn’t go down: From a real Twitter SRE

#367

I know most of my comment is going to be redundant, but I need to say it in order to point to it an year or two from now as an "I told you so!" Twitter is going to be here for a long long time. It's not going to suddenly shut down, it's not going to slowly decline, there is not going to be mass abandonment. Elon might be cocky, but at the end of the day he is a successful businessman. He didn't just sent out that loy…

He doesn’t understand shit, you don’t fire that many people before even having as crude of an understanding as was seen on that whiteboard. Also, people working on embedded systems for rockets know jack-shit about a system serving petabytes of data.

Re: Why Twitter didn’t go down: From a real Twitter SRE

#368
post #147
post #95

Earlier quoted context omitted.

Want to bet that’s not unconnected to sacking the people who tried to block bots?

I'm sure it's connected, but not because of the bots. Millions of political moderates are stopping by to see what happens when a large-scale public forum begins to tolerate freedom of speech.

Interestingly, my reports on racist posts are being accepted much more often than they were before, so this is in fact not happening.

The site was also consumed in horrific discourse for the last week (topic being "is making chili for your neighbor racist, ableist and oppressing autistic people?") so all in all it seems to behaving as usual.

Re: Why Twitter didn’t go down: From a real Twitter SRE

#369

Earlier quoted context omitted.

If the 3x headcount increase really did add no value, there are still about 1/3rd profitable employees there now. In fact giant layoffs tend to cut the best people first because they are the ones who feel comfortable walking. The people that are the last to go are the ones who are very entrenched in the organization and who don't estimate their chances outside of it highly, and that's the exact description of who Elo…

I've found the opposite. I almost never see low performing employees fired outside of a mass layoff. In every layoffs I've seen 10x as many people were fired as quit. So you lose a bunch of low performers involuntarily, and a few top performers both voluntarily and involuntarily and that leads to the average quality improve.

Do you think they did the work necessary to tell a high performer from a low one?

The short amount of time they took and methods they reportedly used don't instill much confidence.

Re: Why Twitter didn’t go down: From a real Twitter SRE

#370
post #237

Earlier quoted context omitted.

We had an obituary for Fred Brooks on here just the other day. I'd suggest that his thesis in The Mythical Man-Month conflicts with your comment above (that reduction in staff count for a software project has a good correlation with the ability to maintain it / evolve it / innovate on top of it).

I've never heard of (or thought of) your interpretation of the corollary to Brooke's Law, but removing people from projects until they succeed and are on time seems like a bold strategy.

Happens regularly in the Free Software & Open Source world, where software projects are often forked by a single individual. Obviously enterprise software is on another level, but the principle remains that reduction to a smaller development team (for some period of time) does not necessarily correlate to a threat to viability and has often indeed been a reinjection of vitality.
Post reply on HN