Live data from Hacker News

Production Twitter on one machine? 100Gbps NICs and NVMe are fast

thume.ca

131–140 of 500 posts

Re: Production Twitter on one machine? 100Gbps NICs and NVMe are fast

#131

Earlier quoted context omitted.

Well you gotta have a backup strategy. I'm talking about the primary machine here, I assumed that would be obvious but maybe not. You build your failover strategy into your architecture - there's lots of ways to do it - I use Postgres so I would favor something based around log shipping.

And uptime is important, so you want to have that secondary running and ready, with a proxy in front of everything so you can switch as soon as you detect a failure. That's three hosts, plus your alerting has to be separate too, so that's four. Now, to orchestrate all this, we'll first get out Puppet...

If you're going from one machine to two, and you add an automatically failover mechansim, chances are your load switching mechanism is going to cause more downtime than just running from your single machine, and manually switching on failure (after being paged).

Re: Production Twitter on one machine? 100Gbps NICs and NVMe are fast

#132
post #23

John Carmack tweeted something that made me noodle on this too: >It is amusing to consider how much of the world you could serve something like Twitter to from a single beefy server if it really was just shuffling tweet sized buffers to network offload cards. Smart clients instead of web pages could make a very large difference. [1] Very interesting to see the idea worked out in more detail. [1] https://twitter.com/i…

> just shuffling tweet sized buffers to network offload cards

Except that's not what it is doing at all.

It assembles all the Tweets internally, applies an ML model to produce a finalised response to the user.

Re: Production Twitter on one machine? 100Gbps NICs and NVMe are fast

#133

Earlier quoted context omitted.

> However, it does show how much real compute bloat these companies actually have No, it doesn’t. It’s a fun exercise in approaching Twitter as an academic exercise. It ignores all of the real-world functionality that makes it a business rather than a toy. A lot of complicated businesses are easy to prototype out if you discard all requirements other than the core feature. In the real world, more engineering work oft…

Genuinely asking, why do you think Twitter needs 24 million vcpus to run? This is not apples to apples but Whatsapp is a product that entirely ran on 16 servers at the time of acquisition (1.5 billion users). It really begs the question why Twitter uses so much compute if there are companies that have operated significantly more efficiently. Twitter was unprofitable during acquisition and spent around half their reve…

> This is not apples to apples but Whatsapp is a product that entirely ran on 16 servers at the time of acquisition (1.5 billion users).

- 450m DAUs at the time of facebook acquisition [0]

- Twitter is not just DMs or Group Chat.

> It really begs the question why Twitter uses so much compute if there are companies that have operated significantly more efficiently.

A fair comparision might have been Instagram: While Systrom did run a relatively lean eng org, they never had to monetize and got acquired before they got any bigger than ~50m?

[0] https://www.sequoiacap.com/article/four-numbers-that-explain...

Re: Production Twitter on one machine? 100Gbps NICs and NVMe are fast

#134
post #89

Earlier quoted context omitted.

You don't need k8s for that, teams have been doing that for decades before k8s was ever a thing

Sure, but whatever they built themselves to accomplish this is also complicated. I know because I have built such systems (and replaced them by k8s).

Hah, exactly. It's not that you can't accomplish all the same things as k8s with your own bash scripts - it's that k8s exists to replace all your custom bash scripts!

Re: Production Twitter on one machine? 100Gbps NICs and NVMe are fast

#135

Earlier quoted context omitted.

> However, it does show how much real compute bloat these companies actually have No, it doesn’t. It’s a fun exercise in approaching Twitter as an academic exercise. It ignores all of the real-world functionality that makes it a business rather than a toy. A lot of complicated businesses are easy to prototype out if you discard all requirements other than the core feature. In the real world, more engineering work oft…

Genuinely asking, why do you think Twitter needs 24 million vcpus to run? This is not apples to apples but Whatsapp is a product that entirely ran on 16 servers at the time of acquisition (1.5 billion users). It really begs the question why Twitter uses so much compute if there are companies that have operated significantly more efficiently. Twitter was unprofitable during acquisition and spent around half their reve…

Whatsapp used Erlang.

Re: Production Twitter on one machine? 100Gbps NICs and NVMe are fast

#136

Earlier quoted context omitted.

My friend mentioned this just before I published and I think that probably is the fastest largest thing you can get which would in some sense count as one machine. I haven't looked into it, but I wouldn't be surprised if they could get around the trickiest constraint, which is how many hard drives you can plug in to a non-mainframe machine for historical image storage. Definitely more expensive than just networking a…

Incidentally, a lot of people have argued that the massive datacenters used by e.g. AWS are effectively single large ("warehouse-scale") computers. In a way, it seems that the mainframe has been reinvented.

I wouldn't really agree with this since those machines don't share address spaces or directly attached busses. Better to say it's a warehouse-scale "service" provided by many machines which are aggregated in various ways.

Re: Production Twitter on one machine? 100Gbps NICs and NVMe are fast

#137
post #93

Earlier quoted context omitted.

It's a neat thought exercise, but wrong for so many reasons (there are probably like 100s). Some jump out: spam/abuse detection, ad relevance, open graph web previews, promoted tweets that don't appear in author timelines, blocks/mutes, etc. This program is what people think Twitter is, but there's a lot more to it. I think every big internet service uses user-space networking where required, so that part isn't new.

I think I'm pretty careful to say that this is a simplified version of Twitter. Of the features you list: - spam detection: I agree this is a reasonably core feature and a good point. I think you could fit something here but you'd have to architect your entire spam detection approach around being able to fit, which is a pretty tricky constraint and probably would make it perform worse than a less constrained solution…

"web previews: I'd do this by making it the client's responsibility."

Actually a good example of how difficult the problem is. A very common attack is to switch a bit.ly link or something like that to a malicious destination. You would also DoS the hosts... as the Mastodon folks are discovering (https://www.jwz.org/blog/2022/11/mastodon-stampede/)

For blocks/mutes, you have to account for retweets and quotes, it's just not a fun problem.

Shipping the product is much more difficult that what's in your post. It's not realistic at all, but it is fun to think about.

Re: Production Twitter on one machine? 100Gbps NICs and NVMe are fast

#139

Earlier quoted context omitted.

> spending more than this for dozens of kubernetes clusters on AWS before signing a single customer Yup. Cloud is 21st century Nickel & Diming. Sure it sounds cheap, everything is priced in small sounding cents per unit. But then it very quickly becomes a compounding vicious circle ... a dozen different cloud tools, each charged at cents per unit, those units often being measured in increments of hours....next thing…

> You can get quite a lot of bang for your buck in a 1/4 rack with today's kit. Until very recently, while money was still very cheap, the time overhead it would take to manage this just was not worth the cost savings. Even with the market falling out from under VC, I think it still is a good tradeoff for many shops.

> Until very recently, while money was still very cheap, the time overhead it would take to manage this just was not worth the cost savings.

You can also rent a whole server. There's not much difference in time in managing a VM in a cloud or a whole server you rent from someone. Depending on the vendor, maybe some more setup time, since low end hosts don't usually have great setup workflows, so maybe you need to fiddle with the ipmi console once or twice to get it started, but if you go with a higher tier provider, you can fully automate everything if that floats your boat. It's just bare metal rather than a VM, and typically much lower cost for sustained usage (if you're really scaling up signfigantly and down throughout the day, cloud costs can work out less, although some vendors offer bare metal by the hour, too)

Re: Production Twitter on one machine? 100Gbps NICs and NVMe are fast

#140

How much bandwidth does Twitter use for images and videos? Less than 1.4Tb/s globally? If so, we could probably fit that onto a second machine. We can currently serve over 700Gb/s from a dual-socket Milan based server[1]. I'm still waiting for hardware, but assuming there are no new bottlenecks, that should directly scale up to 1.4Tb/s with Genoa and ConnectX-7, given the IO pathways are all at least twice the bandwi…

It is way more than 1.4TBs a second globally.

I wonder how much is api traffic and how much is assets & images.
Post reply on HN