Live data from Hacker News

macOS 26.2 enables fast AI clusters with RDMA over Thunderbolt

developer.apple.com

271–280 of 304 posts

Re: macOS 26.2 enables fast AI clusters with RDMA over Thunderbolt

#271
post #252

Earlier quoted context omitted.

I did the same, then put in 14 3090's. It's a little bit power hungry but fairly impressive performance wise. The hardest parts are power distribution and riser cards but I found good solutions for both.

You get occasional accounts of 3090 home-superscalers whereas they would put up eight, ten, fourteen cards. I normally attribute this to obsessive-compulsive behaviour. What kind of motherboard you ended up using and what's the bi-directional bandwidth you're seeing? Something tells me you're not using EPYC 9005's with up to 256x PCIe 5.0 lanes per socket or something... Also: I find it hard to believe the "performan…

I love your skepsis of what I consider to be a fairly normal project, this is not to brag, simply to document.

And I'm way above 3 kW, more likely 5000 to 5500 with the GPUs running as high as I'll let them, or thereabouts, but I only have one power meter and it maxes out at 2500 watts or so. This is using two Xeons in a very high end but slightly older motherboard. When it runs the space that it is in becomes hot enough that even in the winter I have to use forced air from outside otherwise it will die.

As for electricity costs, I have 50 solar panels and on a good day they more than offset the electricity use, at 2 pm (solar noon here) I'd still be pushing 8 KW extra back into the grid. This obviously does not work out so favorably in the winter.

Building a system like this isn't very hard, it is just a lot of money for a private individual but I can afford it, I think this build is a bit under $10K, so a fraction of what you'd pay for a commercial solution but obviously far less polished and still less performant. But it is a lot of bang for the buck and I'd much rather have this rig at $10K than the first commercial solution available at a multiple of this.

I wrote a bit about power efficiency in the run-up to this build when I only had two GPUs to play with:

https://jacquesmattheij.com/llama-energy-efficiency/

My main issue with the system is that it is physically fragile, I can't transport it at all, you basically have to take it apart and then move the parts and re-assemble it on the other side. It's just too heavy and the power distribution is messy so you end up with a lot of loose wires and power supplies. I could make a complete enclosure for everything but this machine is not running permanently and when I need the space for other things I just take it apart, store the GPUs in their original boxes until the next home-run AI project. Putting it all together is about 2 hours of work. We call it Frankie, on account of how it looks.

edit: one more note, the noise it makes is absolutely incredible and I would not recommend running something like this in your house unless you are (1) crazy or (2) have a separate garage where you can install it.

Re: macOS 26.2 enables fast AI clusters with RDMA over Thunderbolt

#272
post #89

Earlier quoted context omitted.

What's the math on the $50k nvidia cluster? My understanding these things cost ~$8k and you can at least get 5 for $40k, that's around half a tb. That being said, for inference mac still remain the best, and the M5 Ultra will even be a better value with its better PP.

GPUs: 4x NVIDIA RTX 6000 Blackwell (96GB VRAM each) • Cost: 4 × $9,000 = $36,000 • CPU: AMD Ryzen Threadripper PRO 7995WX (96-Core) • Cost: $10,000 • Motherboard: WRX90 Chipset (supports 7x PCIe Gen5 slots) • Cost: $1,200 • RAM: 512GB DDR5 ECC Registered • Cost: $2,000 • Chassis & Power: Supermicro or specialized Workstation case + 2x 1600W PSUs. • Cost: $1,500 • Total Cost: ~$50,700 It’s a bit maximalist, but if you…

This is basically a tinybox pro?

Re: macOS 26.2 enables fast AI clusters with RDMA over Thunderbolt

#273

Earlier quoted context omitted.

Outside of YouTube influencers, I doubt many home users are buying a 512G RAM Mac Studio.

I'm neither and have 2. 24/7 async inference against github issues. Free. (once you buy the macs that is)

Interesting. Answering them? Solving them? Looking for ones to solve?

Re: macOS 26.2 enables fast AI clusters with RDMA over Thunderbolt

#274
post #152
post #144

Earlier quoted context omitted.

"... Thunderbolt is great for peripherals, but as a semi-permanent interconnect, I have worries over the port's physical stability ..." Thunderbolt as a server interconnect displeases me aesthetically but my conclusion is the opposite of yours: If the systems are locked into place as servers in a rack the movements and stresses on the cable are much lower than when it is used as a peripheral interconnect for a deskto…

This is a semi-solved problem e.g. https://www.sonnetstore.com/products/thunderlok-a Apple’s chassis do not support it. But conceptually that’s not a Thunderbolt problem, it’s an Apple problem. You could probably drill into the Mac Studio chassis to create mount points.

You could also epoxy it.

Re: macOS 26.2 enables fast AI clusters with RDMA over Thunderbolt

#275

Anyone found any APIs related to this? I'd have some other uses for RDMA between Macs.

I found some useful clues here. Looks like it uses the regular InfiniBand RDMA APIs.

https://github.com/Anemll/mlx-rdma/commit/a901dbd3f9eeefc628...

Re: macOS 26.2 enables fast AI clusters with RDMA over Thunderbolt

#277

Earlier quoted context omitted.

I’m not particularly interested in training models, but it would be nice to have eGPUs again. When Apple Silicon came out, support for them dried up. I sold my old BlackMagic eGPU. That said, the need for them also faded. The new chips have performance every bit as good as the eGPU-enhanced Intel chips.

eGPU with an Apple accelerator with a bunch or RAM and GPU cores could be really interesting honestly. I’m pretty sure they are capable of designing something very competitive especially in terms of performance per watt.

Really, that’s a place for the MacPro: slide in SoC with ram modules / blades. Put 4, 8, 16 Ultra chips in one machine.

Re: macOS 26.2 enables fast AI clusters with RDMA over Thunderbolt

#278

Earlier quoted context omitted.

Essentially, the Grace CPU is a memory and IO expander that happens to have a bunch of ARM CPU cores filling in the interior of the die, while the perimeter is all PHYs for LPDDR5 and NVLink and PCIe.

> have a bunch of ARM CPU cores filling in the interior of the die The main OS needs to run somewhere. At least for now.

Sure, but 72x Neoverse V3 (approximately Cortex X3) is a choice that seems more driven by convenience than by any real need for an AI server to have tons of somewhat slow CPU cores.

Re: macOS 26.2 enables fast AI clusters with RDMA over Thunderbolt

#279

Earlier quoted context omitted.

What pcie version are you running? Normally I would not mention one of these, but you have already invested in all the cards, and it could free up some space if any of your lanes being used now are 3.0. If you can afford the 16 (pcie 3) lanes, you could get a PLX ("PCIe Gen3 PLX Packet switch X16 - x8x8x8x8" on ebay for like $300) and get 4 of your cards up to x8.

All are PCIe 3.0, I wasn't aware of those switches at all, in spite of buying my risers and cables from that source! Unfortunately all of the slots on the board are x8, there are no x16 slots at all. So that switch would probably work but I wonder how big the benefit would be: you will probably see effectively an x4 -> (x4 / x8) -> (x8 / x8) -> (x8 / x8) -> (x8 / x4) -> x4 pipeline, and then on to the next set of fou…

If you end up trying it please share your findings!

I've basically been putting this kind of gear in my cart, and then deciding I dont want to manage more than the 2 3090s, 4090 and a5000 I have now, then I take the PLX out of my cart.

Seeing you have the cards already it could be a good fit!

Re: macOS 26.2 enables fast AI clusters with RDMA over Thunderbolt

#280

Earlier quoted context omitted.

All are PCIe 3.0, I wasn't aware of those switches at all, in spite of buying my risers and cables from that source! Unfortunately all of the slots on the board are x8, there are no x16 slots at all. So that switch would probably work but I wonder how big the benefit would be: you will probably see effectively an x4 -> (x4 / x8) -> (x8 / x8) -> (x8 / x8) -> (x8 / x4) -> x4 pipeline, and then on to the next set of fou…

If you end up trying it please share your findings! I've basically been putting this kind of gear in my cart, and then deciding I dont want to manage more than the 2 3090s, 4090 and a5000 I have now, then I take the PLX out of my cart. Seeing you have the cards already it could be a good fit!

Yes, it could be. Unfortunately I'm a bit distracted by both paid work and some more urgent stuff but eventually I will get back to it. By then this whole rig might be hopelessly outdated but we've done some fun experiments with it and have kept our confidential data in-house which was the thing that mattered to me.
Post reply on HN