Live data from Hacker News

macOS 26.2 enables fast AI clusters with RDMA over Thunderbolt

developer.apple.com

241–250 of 304 posts

Re: macOS 26.2 enables fast AI clusters with RDMA over Thunderbolt

#242
post #3

I follow the MLX team on Twitter and they sometimes post about using MLX on two or more joined together Macs to run models that need more than 512GB of RAM. A couple of examples: Kimi K2 Thinking (1 trillion parameters): https://x.com/awnihannun/status/1986601104130646266 DeepSeek R1 (671B): https://x.com/awnihannun/status/1881915166922863045 - that one came with setup instructions in a Gist: https://gist.github.com/…

For a bit more context, those posts are using pipeline parallelism. For N machines put the first L/N layers on machine 1, next L/N layers on machine 2, etc. With pipeline parallelism you don't get a speedup over one machine - it just buys you the ability to use larger models than you can fit on a single machine. The release in Tahoe 26.2 will enable us to do fast tensor parallelism in MLX. Each layer of the model is…

Exo-Labs is an open source project that allows this too, pipeline parallelism I mean not the latter, and it's device agnostic meaning you can daisy-chain anything you have that has memory and the implementation will intelligently shard model layers across them, though its slow but scales linearly with concurrent requests.

Exo-Labs: https://github.com/exo-explore/exo

Re: macOS 26.2 enables fast AI clusters with RDMA over Thunderbolt

#243

dang I wish I could share md tables. Here’s a text edition: For $50k the inference hardware market forces a trade-off between capacity and throughput: * Apple M3 Ultra Cluster ($50k): Maximizes capacity (3TB). It is the only option in this price class capable of running 3T+ parameter models (e.g., Kimi k2), albeit at low speeds (~15 t/s). * NVIDIA RTX 6000 Workstation ($50k): Maximizes throughput (>80 t/s). It is sup…

what about a GB300 workstation with 784GB unified mem?

Re: macOS 26.2 enables fast AI clusters with RDMA over Thunderbolt

#244

Earlier quoted context omitted.

You can keep scaling down! I spent $2k on an old dual-socket xeon workstation with 768GB of RAM - I can run Deepseek-R1 at ~1-2 tokens/sec.

Nice! What do you use it for?

1-2 tokens/sec is perfectly fine for 'asynchronous' queries, and the open-weight models are pretty close to frontier-quality (maybe a few months behind?). I frequently use it for a variety of research topics, doing feasibility studies for wacky ideas, some prototypy coding tasks. I usually give it a prompt and come back half an hour later to see the results (although the thinking traces are sufficiently entertaining that sometimes it's fun to just read as it comes out). Being able to see the full thinking traces (and pause and alter/correct them if needed) is one of my favorite aspects of being able to run these models locally. The thinking traces are frequently just as or more useful than the final outputs.

Re: macOS 26.2 enables fast AI clusters with RDMA over Thunderbolt

#245

This implies you'd run more than one Mac Studio in a cluster, and I have a few concerns regarding Mac clustering (as someone who's managed a number of tiny clusters, with various hardware): 1. The power button is in an awkward location, meaning rackmounting them (either 10" or 19" rack) is a bit cumbersome (at best) 2. Thunderbolt is great for peripherals, but as a semi-permanent interconnect, I have worries over the…

VNC over SSH tunneling always worked well for me before I had Apple Remote Desktop available, though I don't recall if I ever initiated a connection attempt from anything other than macOS...

erase-install can be run non-interactively when the correct arguments are used. I've only ever used it with an MDM in play so YMMV:

https://github.com/grahampugh/erase-install

Re: macOS 26.2 enables fast AI clusters with RDMA over Thunderbolt

#246
post #64

Earlier quoted context omitted.

That doesn't answer the question, which was how to get a high-speed interconnect between a Mac and a DGX Spark. The most likely solution would be a Thunderbolt PCIe enclosure and a 100Gb+ NIC, and passive DAC cables. The tricky part would be macOS drivers for said NIC.

You’re right I misunderstood. I’m not sure if it would be of much utility because this would presumably be for tensor parallel workloads. In that case you want the ranks in your cluster to be uniform or else everything will be forced to run at the speed of the slowest rank. You could run pipeline parallel but not sure it’d be that much better than what we already have.

It was about this use case:

https://blog.exolabs.net/nvidia-dgx-spark/

Re: macOS 26.2 enables fast AI clusters with RDMA over Thunderbolt

#247

Earlier quoted context omitted.

Exactly: The AI appliance market. A new kind of home or small-business server.

I’m expecting Apple to release a new Mac Pro in the next couple years who’s main marketing angle is exactly this

It’s really the only common reason to buy a machine that big these days. I could see a Mac Pro with a huge GPU and up to a terabyte of RAM.

I guess there are other kinds of scientific simulation, very large dev work, and etc., but those things are quite a bit more niche.

Re: macOS 26.2 enables fast AI clusters with RDMA over Thunderbolt

#248

Earlier quoted context omitted.

I haven’t looked yet but I might be a candidate for something like this, maybe. I’m RAM constrained and, to a lesser extent, CPU constrained. It would be nice to offload some of that. That said, I don’t think I would buy a cluster of Macs for that. I’d probably buy a machine that can take a GPU.

I’m not particularly interested in training models, but it would be nice to have eGPUs again. When Apple Silicon came out, support for them dried up. I sold my old BlackMagic eGPU. That said, the need for them also faded. The new chips have performance every bit as good as the eGPU-enhanced Intel chips.

eGPU with an Apple accelerator with a bunch or RAM and GPU cores could be really interesting honestly. I’m pretty sure they are capable of designing something very competitive especially in terms of performance per watt.

Re: macOS 26.2 enables fast AI clusters with RDMA over Thunderbolt

#249

Earlier quoted context omitted.

Exactly: The AI appliance market. A new kind of home or small-business server.

I’m expecting Apple to release a new Mac Pro in the next couple years who’s main marketing angle is exactly this

I fear they no longer care about the workstation market, even the folks at ATP Podcast are at the verge of accepting it.

Re: macOS 26.2 enables fast AI clusters with RDMA over Thunderbolt

#250
post #233

Earlier quoted context omitted.

> You don't need to be a genius or a billionaire to realize that when most of the global supply of a product becomes unavailable the remaining supply gets more expensive. Yes. Absolutely correct if you are talking about the short term. I was talking about the long term, and said that. If you are so certain would you take this bet: any odds, any amount that within 1 month I can buy 32gb of new retail DDR5 in the US fo…

What is your estimate for when memory prices will decrease? I agree that we've seen similar fluctuations in the past and the price of compute trends down in the long-term. This could be a bubble, which it likely is, in which case prices should return to baseline eventually. The political climate is extremely challenging at this time though so things could take longer to stabilize. Do you think we're in this ride for…

I can’t be more clear: specificity around predicting the future is close to impossible. There are 9 figure bets on both sides of the RAM issue, and strategic national concerns. I say that prices will go down at some point in the future for reasons highlighted already, but I have no clue when. Keep in mind what I myself have said about human ability to predict the future. You would be a fool to believe anyone’s specific estimates.

Maybe the AI money train stops after Christmas. The entire economy is fucked, but RAM is cheap.

Maybe we unlock AGI and the price sky rockets further before factories can get built.

There are just too many variables.

The real test is if someone had seen this coming, they would have made massive absurd investment returns just by buying up stock and storing it for a few months. Anyone who didn’t take advantage of that opportunity has proved that they had no real confidence in their ability to predict the future price of RAM. RAM inventory might have been one of the highest return investments possible this year. Where are all the RAM whales in Lambos who saw this coming?

As a corollary: we can say that unless you have some skin in the game and have invested a significant amount of your wealth in RAM chips, then you don’t know which way the price is going or when.

Extending that even further: people complaining about RAM prices being so high, and moaning that they bought less RAM because of it are actually signaling through action that they think that prices will go down or have leveled off. Anyone who believes that sticks of DDR5 RAM will continue the trend should be cleaning out Amazon, Best Buy and Newegg since the price will never be lower than today.

The distinct lack of serious people saying “I told ya so” with receipts, combined with the lack of people hoarding RAM to sell later is a good indirect signal that no one knows what is happening in the near term.

Post reply on HN