Live data from Hacker News

macOS 26.2 enables fast AI clusters with RDMA over Thunderbolt

developer.apple.com

251–260 of 304 posts

Re: macOS 26.2 enables fast AI clusters with RDMA over Thunderbolt

#251

Earlier quoted context omitted.

Are the inference providers profitable yet? Might be nice to be ready for the day when we see the real price of their services.

Isn't it then even better to enjoy cheap inference thanks to techbro philanthropy while it lasts? You can always buy the hardware once the free money runs out.

Probably depends on what you are interested in. IMO, setting up local programs is more fun anyway. Plus, any project I’d do with LLMs would just be for fun and learning at this point, so I figure it is better to learn skills that will be useful in the long run.

Re: macOS 26.2 enables fast AI clusters with RDMA over Thunderbolt

#252

Earlier quoted context omitted.

You can keep scaling down! I spent $2k on an old dual-socket xeon workstation with 768GB of RAM - I can run Deepseek-R1 at ~1-2 tokens/sec.

I did the same, then put in 14 3090's. It's a little bit power hungry but fairly impressive performance wise. The hardest parts are power distribution and riser cards but I found good solutions for both.

You get occasional accounts of 3090 home-superscalers whereas they would put up eight, ten, fourteen cards. I normally attribute this to obsessive-compulsive behaviour. What kind of motherboard you ended up using and what's the bi-directional bandwidth you're seeing? Something tells me you're not using EPYC 9005's with up to 256x PCIe 5.0 lanes per socket or something... Also: I find it hard to believe the "performance" claims, when your rig is pulling 3 kW from the wall (assuming undervolting at 200W per card?) The electricity costs alone would surely make this intractable, i.e. the same as running six washing machines all at once.

Re: macOS 26.2 enables fast AI clusters with RDMA over Thunderbolt

#253
post #234

Earlier quoted context omitted.

I think 14 3090's are more than a little power hungry!

to the point that I had to pull an extra circuit... but tri phase so good to go even if I would like to go bigger. I've limited power consumption to what I consider the optimum, each card will draw ~275 Watts (you can very nicely configure this on a per-card basis). The server itself also uses some for the motherboard, the whole rig is powered from 4 1600W supplies, the gpus are divided 5/5/4 and the mother board is…

What pcie version are you running? Normally I would not mention one of these, but you have already invested in all the cards, and it could free up some space if any of your lanes being used now are 3.0.

If you can afford the 16 (pcie 3) lanes, you could get a PLX ("PCIe Gen3 PLX Packet switch X16 - x8x8x8x8" on ebay for like $300) and get 4 of your cards up to x8.

Re: macOS 26.2 enables fast AI clusters with RDMA over Thunderbolt

#254
post #193

Earlier quoted context omitted.

Not Apple’s ram.

RAM prices have exploded enough that Apple's RAM is now no longer a bad deal. At least until their next price hikes. We're going back to the "consumer PCs have 8GB of RAM era" thanks to the AI bubble.

Funny, considering Macbooks finally started shipping at 16 GB due to Apple Intelligence.

Re: macOS 26.2 enables fast AI clusters with RDMA over Thunderbolt

#255
post #127

Earlier quoted context omitted.

Just keep going! 2TB of swap disk for 0.0000001 t/sec

Hang on, starting benchmarks on my Raspberry Pi.

On a lark a friend setup Ollama on a 8GB Raspberry Pi with one of the smaller models. It worked by it was very slow. IIRC it did 1 token/second.

Re: macOS 26.2 enables fast AI clusters with RDMA over Thunderbolt

#258

Earlier quoted context omitted.

Outside of YouTube influencers, I doubt many home users are buying a 512G RAM Mac Studio.

Of course they're not. Everybody is waiting for next generation that will run LLMs faster to start buying.

Every generation runs LLMs faster than the previous one.

Re: macOS 26.2 enables fast AI clusters with RDMA over Thunderbolt

#259

Earlier quoted context omitted.

Seems like it could be a thing. Also, I’m curious and in case anyone that knows reads this comment: Apple say they can’t get the performance they want out of discreet GPUs. Fair enough. But yet nVidia becomes the most valuable company in the world selling GPUs. So… Now I get that Apples use case is essentially sealed consumer devices built with power consumption and performance tradeoffs in mind. But could Apple use…

There’s been rumors of Apple working on M-chips that have the GPU and CPU as discrete chiplets. The original rumor said this would happen with the M5 Pro, so it’s potentially on the roadmap. Theoretically they could farm out the GPU to another company but it seems like they’re set on owning all of the hardware designs.

TSMC has a new tech that allows seamless integration of mini chiplets, i.e. you can add as many CPU/GPU cores in mini chiplets as you wish and glue them seamlessly together, at least in theory. The rumor is that TSMC had some issues with it which is why M5P and M5M are delayed.

Re: macOS 26.2 enables fast AI clusters with RDMA over Thunderbolt

#260

Earlier quoted context omitted.

the GH/GB compute has LPDDR5X - a single or dual GPU shares 480GB, depending if it's GH or GB, in addition to the HBM memory, with NVLink C2C - it's not bad!

Essentially, the Grace CPU is a memory and IO expander that happens to have a bunch of ARM CPU cores filling in the interior of the die, while the perimeter is all PHYs for LPDDR5 and NVLink and PCIe.

> have a bunch of ARM CPU cores filling in the interior of the die

The main OS needs to run somewhere. At least for now.

Post reply on HN