Live data from Hacker News

macOS 26.2 enables fast AI clusters with RDMA over Thunderbolt

developer.apple.com

31–40 of 304 posts

Re: macOS 26.2 enables fast AI clusters with RDMA over Thunderbolt

#31

Earlier quoted context omitted.

FLOPS are not what matters here.

also cheaper memory bandwidth. where are you claiming that M5 wins?

I'm not sure where else you can get a half TB of 800GB/s memory for < $10k. (Though that's the M3 Ultra, don't know about the M5). Is there something competitive in the nvidia ecosystem?

Re: macOS 26.2 enables fast AI clusters with RDMA over Thunderbolt

#32
post #11
post #9

It would be incredibly ironic if, with Apple's relatively stable supply chain relative to the chaos of the RAM market these days (projected to last for years), Apple compute became known as a cost-effective way to build medium-sized clusters for inference.

It’s gonna suck if all the good Macs get gobbled up by commercial users.

it's not like regular people can afford this kind of Apple machine anyway.

Re: macOS 26.2 enables fast AI clusters with RDMA over Thunderbolt

#33

This implies you'd run more than one Mac Studio in a cluster, and I have a few concerns regarding Mac clustering (as someone who's managed a number of tiny clusters, with various hardware): 1. The power button is in an awkward location, meaning rackmounting them (either 10" or 19" rack) is a bit cumbersome (at best) 2. Thunderbolt is great for peripherals, but as a semi-permanent interconnect, I have worries over the…

It’s been terrible for years/forever. Even Xserves didn’t really meet the needs of a professional data centre. And it’s got worse as a server OS because it’s not a core focus. Don’t understand why anyone tries to bother - apart from this MLX use case or as a ProRes render farm.

Re: macOS 26.2 enables fast AI clusters with RDMA over Thunderbolt

#34
post #7

Very cool. It requires a fully-connected mesh so the scaling limit here would seem to be 6 Mac Studio M3 Ultra, up to 3TB of unified memory to work with.

I'm sure someone will figure out how to make thunderbolt switch/router

I don't believe the standard supports such a thing. But I wonder if TB6 will.

Re: macOS 26.2 enables fast AI clusters with RDMA over Thunderbolt

#35

Earlier quoted context omitted.

also cheaper memory bandwidth. where are you claiming that M5 wins?

I'm not sure where else you can get a half TB of 800GB/s memory for < $10k. (Though that's the M3 Ultra, don't know about the M5). Is there something competitive in the nvidia ecosystem?

I wasn't aware that M3 Ultra offered a half terabyte of unified memory, but an RTX5090 has double that bandwidth and that's before we even get into B200 (~8TB/s).

Re: macOS 26.2 enables fast AI clusters with RDMA over Thunderbolt

#37
post #9

It would be incredibly ironic if, with Apple's relatively stable supply chain relative to the chaos of the RAM market these days (projected to last for years), Apple compute became known as a cost-effective way to build medium-sized clusters for inference.

It already is depending on your needs.

Re: macOS 26.2 enables fast AI clusters with RDMA over Thunderbolt

#40
post #30
post #3

I follow the MLX team on Twitter and they sometimes post about using MLX on two or more joined together Macs to run models that need more than 512GB of RAM. A couple of examples: Kimi K2 Thinking (1 trillion parameters): https://x.com/awnihannun/status/1986601104130646266 DeepSeek R1 (671B): https://x.com/awnihannun/status/1881915166922863045 - that one came with setup instructions in a Gist: https://gist.github.com/…

I’m hoping this isn’t as attractive as it sounds for non-hobbyists because the performance won’t scale well to parallel workloads or even context processing, where parallelism can be better used. Hopefully this makes it really nice for people that want the experiment with LLMs and have a local model but means well funded companies won’t have any reason to grab them all vs GPUs.

I haven’t looked yet but I might be a candidate for something like this, maybe. I’m RAM constrained and, to a lesser extent, CPU constrained. It would be nice to offload some of that. That said, I don’t think I would buy a cluster of Macs for that. I’d probably buy a machine that can take a GPU.
Post reply on HN