Live data from Hacker News

Why Twitter didn’t go down: From a real Twitter SRE

matthewtejo.substack.com

11–20 of 1001 posts

Re: Why Twitter didn’t go down: From a real Twitter SRE

#13
post #5

Excellent article. However!, software, like everything, is subject to laws of physics. Entropy always wins in the end. No matter how good the original engineering and planning, without maintenance it will all fall apart soon enough.

Right. It’s not like if all the engineers walk out it’s suddenly going to fall over. It’s that when an issue comes up there may not be the expertise on hand to fix it. So the remaining twitter engineers should expect a rough time over the coming months.

There must be some plan for this, though, right? If it were me and I were trying to do more with less I'd plan to cut some features (for instance, are Spaces that critical to the operation? Because it seems like running them would be demanding) and try to migrate things to managed services to reduce the operational load.

Re: Why Twitter didn’t go down: From a real Twitter SRE

#14

Very interesting article. But the constant misuse of “to” for “too” kept throwing me out of reading mode. I wonder why my brain does that instead of reading past it since I know what is actually meant.

Interesting, I didn't notice despite reading the article start too finish. The brain is so mysterious!

Re: Why Twitter didn’t go down: From a real Twitter SRE

#15

Excellent article. However!, software, like everything, is subject to laws of physics. Entropy always wins in the end. No matter how good the original engineering and planning, without maintenance it will all fall apart soon enough.

> Entropy always wins in the end. No matter how good the original engineering and planning, without maintenance it will all fall apart soon enough.

This seems to be more true of Mastodon than Twitter.

I can't imagine any self hosted Mastodon instance staying up longer than twitter.

Re: Why Twitter didn’t go down: From a real Twitter SRE

#16
post #6

I don't think it's that crazy. Private Equity and Wall Street doesn't have pressure for resilience, it has pressure for churn. As long as revenue and spending... exists... things are fine. Ideally they are numbers that grow. This doesn't lend itself to clearing out a development backlog, or engineers doing the most important things, it lends itself to rapid iteration and justifications for the iterations. So now that…

The problem is that it’s run by someone who overpaid badly and needs to come up with significantly more money to pay for the debt he saddled the company with, and then his actions seriously disrupted ad revenue. That puts him more in the PE playbook of cutting costs as deeply as possible even at the extent of long-term growth. Unlike his other companies there isn’t strong government support to drive business for Twitter.

Re: Why Twitter didn’t go down: From a real Twitter SRE

#17
post #8

I wonder how big Twitter's infra costs are (capex). Based on the article, they need to build capacity > 2x peak loaf to tolerate one of the DCs failing. But if they instead had 4 and still only tolerated 1 failure, they'd only need capacity > 1.33x peak load. This doesn't account for the extra overhead associated with extra DCs, but it seems like there's opportunities for major effeciency wins.

Well Elon is planning on shutting down one of their three data centers, the Sacramento data center - so they must have a lot of extra capacity.

https://www.datacenterdynamics.com/en/news/report-elon-musk-...

Re: Why Twitter didn’t go down: From a real Twitter SRE

#18
post #5

Earlier quoted context omitted.

Right. It’s not like if all the engineers walk out it’s suddenly going to fall over. It’s that when an issue comes up there may not be the expertise on hand to fix it. So the remaining twitter engineers should expect a rough time over the coming months.

There must be some plan for this, though, right? If it were me and I were trying to do more with less I'd plan to cut some features (for instance, are Spaces that critical to the operation? Because it seems like running them would be demanding) and try to migrate things to managed services to reduce the operational load.

> There must be some plan for this, though, right?

To be this young and carefree

Re: Why Twitter didn’t go down: From a real Twitter SRE

#19

Excellent article. However!, software, like everything, is subject to laws of physics. Entropy always wins in the end. No matter how good the original engineering and planning, without maintenance it will all fall apart soon enough.

software, like everything, is subject to laws of physics

I disagree; math would be a closer analogy. And indeed, arithmetic still works like it did a millenia ago. Closer to the present, I have binaries from the late 80s that still work today (and I use them semi-regularly.)

Indeed, much of the impetus of the software industry seems to be to propagate the illusion that software somehow needs constant "maintenance" and change. For the preservation of their own self-interests, of course; much like the company that makes physical objects too robust and runs out of customers, planned obsolescence and the desire to change things and justify it so they can be paid to do something are still there.

It's possible to make things which last. Unfortunately, much of the time, other economic considerations preclude that.

Re: Why Twitter didn’t go down: From a real Twitter SRE

#20
post #5

Earlier quoted context omitted.

Right. It’s not like if all the engineers walk out it’s suddenly going to fall over. It’s that when an issue comes up there may not be the expertise on hand to fix it. So the remaining twitter engineers should expect a rough time over the coming months.

There must be some plan for this, though, right? If it were me and I were trying to do more with less I'd plan to cut some features (for instance, are Spaces that critical to the operation? Because it seems like running them would be demanding) and try to migrate things to managed services to reduce the operational load.

I don't have a DR plan for "mad billionaire buys the company", no, but I do agree with the OOM reaper logic - I start cutting the services that use the most resources that are least necessary.
Post reply on HN