Live data from Hacker News

Why Twitter didn’t go down: From a real Twitter SRE

matthewtejo.substack.com

981–990 of 1001 posts

Re: Why Twitter didn’t go down: From a real Twitter SRE

#981

Earlier quoted context omitted.

It would be funny for Musk to share privately with friends. It is a deeply strange and inappropriate thing for the CEO of a billion dollar company to post publicly. Something is deeply wrong with Musk's decision making process.

Is it strange and inappropriate if most people think it is a funny joke and laugh it off? Who goes out of their way to look at Elon posts just to be offended?

You don't have to go out of your way to see it. He posted it publicly and it was widely posted on reddit because it was such a odd thing to post.

Re: Why Twitter didn’t go down: From a real Twitter SRE

#982

Earlier quoted context omitted.

> As long as nothing changes, and you don't run out of disk space (from logs for example), things stay working pretty much just fine. > ... > There are other things that can bring a site down, like security issues, or bugs triggered by unusual states, too much traffic, etc. But generally speaking those things are rare and don't bring down an entire site. Aren't these changes inevitable, though? There is no such thing…

Software never goes stale, it's the environment around it which stales. Something from the 70s works perfectly fine, except it can't run on anything bare any longer, and the hard drives etc. have all long since failed or their PSU capacitors have blown....so Twitter will absolutely rot, how fast depends on several factors. I personally suspect the infrastructure used to build Twitter will rot faster than Twitter itse…

Something from the 70's was not connected to the internet where millions of people are using it every day and finding every single edge case, or they are trying to break into it to steal valuable data. It was definitely not beholden to the same government regulations as a social media site running in the 21st century.

Re: Why Twitter didn’t go down: From a real Twitter SRE

#983

Earlier quoted context omitted.

It's already breaking down. HW/SW aside, a lot of the services had teams monitoring them, which don't exist anymore. Some of the side effects of that: - Trends are not working correctly. - Copyright reporting is not working. - Appealing flagged tweets isn't being responded to. It's only a matter of time before these get abused with no one to fix them.

> - Trends are not working correctly. "Trends" were subject to Twitter's "trends blacklist" before; something that leaked a few years ago. Maybe they're working correctly now that they're unencumbered. Can you describe how they're not working now? > - Appealing flagged tweets isn't being responded to. Mine was responded to in a handful of hours. Much faster than I was expecting. > It's only a matter of time before th…

> Can you describe how they're not working now?

The team would moderate junk trends. So some of the trends are literally just random words being spouted. Similar to early day twitter.

> Mine was responded to in a handful of hours.

Recently? That's impressive when the related teams are gone.

> "the walls are closing in"

I am not sure what a YT video about Trump has to do with my comment.

I refer you "in comments": https://news.ycombinator.com/newsguidelines.html

Re: Why Twitter didn’t go down: From a real Twitter SRE

#984

Earlier quoted context omitted.

Twitter had 7,500 employees. most of the roles you mention (security engineers, code reviews, data architecture, HR, internal audit teams, content moderators, scrum master) are not bloat. So the question is what are the other 7000 people doing?

They had over a thousand moderators. So maybe your estimate of how many people is required is a bit off.

Then the question remains what the other 5000 people were doing.

Re: Why Twitter didn’t go down: From a real Twitter SRE

#985

Earlier quoted context omitted.

> the article might actually persuade me that it was all bloat First of all, how does it persuade you of that? The article touches a really small (though incredibly important for up-time) subject. Secondly, in any large company, the majority is 'bloat'. It's security engineers, code reviews, data architecture, HR, internal audit teams, content moderators, ccrum masters and I can keep going. In a start-up many of thes…

Twitter had 7,500 employees. most of the roles you mention (security engineers, code reviews, data architecture, HR, internal audit teams, content moderators, scrum master) are not bloat. So the question is what are the other 7000 people doing?

the answer is - not much.

the power law applies to any big organization. 20% of the people do 80% of the work, whilst 80% of the people are just there for "support".

whatsapp was run by a team of like 20 people or something when they got acquired for $20 billion. for a simple software product, you don't really need that many people. in fact, more people often means bad software. you just need a small group of very talented engineers to run the product and add new features when necessary.

big (and especially public) companies often times need to hire a lot, just to look like a real company.

now that twitter is private, elon has no responsibility to public investors and can focus less on looking like a real company and more on doing what needs to be done to cut bloat/costs and improve product

Re: Why Twitter didn’t go down: From a real Twitter SRE

#986
post #24

The most helpful thing to reflect on in these Twitter operational discussions is the difference between homeostasis and evolution. You can get rid of 80% of the work force and the existing homeostasis systems will keep things running smoothly despite known day-to-day chaos. Where you’re really going to run into trouble is inventing responses to novel chaos and gradually changing times.

I think the opposite is true.

The bigger a ship is, the slower it is to turn.

IBM is a "tech" company that employs 282,000 employees, and when was the last time they invented something? I don't remember the last time I heard IBM in the news about something they made.

The bigger the company, you often times find less innovation and more administration & bureaucracy.

The reason startups can survive is because of its small size that makes it very flexible and adaptable to chaos and change, that gives it the edge over bigger companies.

Re: Why Twitter didn’t go down: From a real Twitter SRE

#987

Earlier quoted context omitted.

Many people who have worked with Musk have shared similar sentiments in interviews. But it seems that people just refuse to believe any of it. People think that there's no way it's possible for someone to be that deeply technical and be a CEO of multiple companies at the same time. I've talked to people about it and they straight up refuse to believe it saying that it's impossible and that any evidence of him being t…

With a handful of tricks or a patsy in your pocket, it's easy enough to pull things like this off. These are all self-reported encounters, which lends some doubt to them; as I've never seen any public performance of his that suggests he has this exceptional intelligence or that he isn't subject to the same amount of irrational thinking that most humans are. You may be able to do some type of rocket equation in your h…

> but if you constantly promise things that aren't ultimately delivered.. people have good reason to question this narrative.

This is the insane thing to me. He's promised a lot of things, but he has also delivered some pretty huge things. Tesla kicked off the electric car migration and has millions of EVs on the road. SpaceX has reusable first stages on their rockets and are the only private company to send humans to space. Just those two things alone are massive achievements. But people look at some things he's promised but has not yet delivered and that somehow is more important than what he has delivered?

Re: Why Twitter didn’t go down: From a real Twitter SRE

#988

> This left a lot wondering what exactly was going on with all those engineers and made it seem like it was all just bloat. I was partly expecting the rest of the article to explain to me why exactly it wasn't just bloat. But it goes on talking about this 1~3-person cache SRE team that built solid infra automation that's really resilient to both hardware and software failures. If anything, the article might actually…

> the article might actually persuade me that it was all bloat First of all, how does it persuade you of that? The article touches a really small (though incredibly important for up-time) subject. Secondly, in any large company, the majority is 'bloat'. It's security engineers, code reviews, data architecture, HR, internal audit teams, content moderators, ccrum masters and I can keep going. In a start-up many of thes…

I'd argue that growth relies on stability, and if you don't have stability, you'll randomly lose growth.

Re: Why Twitter didn’t go down: From a real Twitter SRE

#989
post #338
post #290

automation is indeed key. IRC networks run with no obvious employees and salarys and been running fine since the dawn of the internet. i see Twitter as a centralised, walled IRC. the # was pinched from IRC if u remember.

An IRC network server can serve what, a few hundred thousands of concurrent users’ text messages on a single PC. Twitter has around 200 million daily user, peaking up to a billion, running on their own server farm. They are not remotely comparable, especially if we note that read-write dbs are notoriously hard to scale.

thank you, i was trying to be provocative to learn more.

otoh, to get an automation going is like climbing mount everest the first time? very hard but makes it easier to scale subsequent attempts?

Re: Why Twitter didn’t go down: From a real Twitter SRE

#990

Earlier quoted context omitted.

I think the opposite. Many softwares at its best when the team was small. Software companies have to hire many people because it needs to report growth to investors, headcount is one of the measurement of growth. It is not necessarily good for the product, actually many times, it hurts the product, but overall it is good for the company, the company will enter new areas, can explore new things. What Twitter is doing…

Scaling back up is really hard though. We had a de facto freeze on hiring (not exactly hiring freeze; more of a headcount cap) just shy of a decade ago to focus on our product. During that time, some of our best recruiters left because they basically had nothing to do anymore. The freeze worked: we got rid of some products that weren't getting traction and were able to improve the products that did have traction. But…

It is not only hard, it also may or may not work. It is the same process Twitter has already went through years ago. I have simplified the issue and talked only about the product. I don't disagree with you.
Post reply on HN