Earlier quoted context omitted.
Also remember that the Mx Ultras have 2-3x the memory bandwidth. Looking at the benchmarks even Strix Halo seems to beat the Spark. Buying a 200 Gbps switch is $10k-$100k+ so don't imagine anyone actually will use the interconnect. The logical thing for Nvidia would be to sell a kit with three machines and cabling, and make it a ring with the dual ports per machine. Helps for some scenarios but not others with the 10…
On another note to remember, you can also ring topology mac studios using TB5 for 120Gbps per link with four such ports, all using cheaply available cable
NVIDIA DGX Spark In-Depth Review: A New Standard for Local AI Inference
81–90 of 100 posts
Re: NVIDIA DGX Spark In-Depth Review: A New Standard for Local AI Inference
#82It isn't that good for local LLM inferencing. It's not designed to be as such. It's designed to be a local dev machine for Nvidia server products. It has the same software and hardware stack as enterprise Nvidia hardware. That's what it is designed for. Wait for M5 series Macs for good value local inferencing. I think the M5 Pro/Max are going to be very good values.
I wish I could run Linux on them (the m5)
Re: NVIDIA DGX Spark In-Depth Review: A New Standard for Local AI Inference
#83Re: NVIDIA DGX Spark In-Depth Review: A New Standard for Local AI Inference
#84Earlier quoted context omitted.
$4,000 is actually extremely competitive. Even for an at-home enthusiast setup this price is not our of reach. I was expecting something far higher, that said, nVidia's MSRP is something of a pipe dream recently so we'll see when it's actually released and the availability. Curious also to see how they may scale together.
A warning to any home consumer throwing money at hardware for AI (fair enough if you have other use cases)... Things are changing rapidly and there is a non insignificant chance that it'll seem like a big waste of money within 12 months.
Re: NVIDIA DGX Spark In-Depth Review: A New Standard for Local AI Inference
#85Earlier quoted context omitted.
Curious to how this compares to running on a Mac.
TTFT on a Mac is terrible and only increases as the context increases, thats why many are selling their M3 Ultra 512GB
https://www.ebay.com/sch/i.html?_nkw=mac+studio+m3+ultra+512...
Re: NVIDIA DGX Spark In-Depth Review: A New Standard for Local AI Inference
#86It isn't that good for local LLM inferencing. It's not designed to be as such. It's designed to be a local dev machine for Nvidia server products. It has the same software and hardware stack as enterprise Nvidia hardware. That's what it is designed for. Wait for M5 series Macs for good value local inferencing. I think the M5 Pro/Max are going to be very good values.
because of possible hardware-accelerated matmul in GPU cores?
Re: NVIDIA DGX Spark In-Depth Review: A New Standard for Local AI Inference
#87It isn't that good for local LLM inferencing. It's not designed to be as such. It's designed to be a local dev machine for Nvidia server products. It has the same software and hardware stack as enterprise Nvidia hardware. That's what it is designed for. Wait for M5 series Macs for good value local inferencing. I think the M5 Pro/Max are going to be very good values.
Re: NVIDIA DGX Spark In-Depth Review: A New Standard for Local AI Inference
#88Re: NVIDIA DGX Spark In-Depth Review: A New Standard for Local AI Inference
#89It isn't that good for local LLM inferencing. It's not designed to be as such. It's designed to be a local dev machine for Nvidia server products. It has the same software and hardware stack as enterprise Nvidia hardware. That's what it is designed for. Wait for M5 series Macs for good value local inferencing. I think the M5 Pro/Max are going to be very good values.
I am still amazed at how many companies buy a ton of DGX boxes and then are surprised that Nvidia does not have any Kubernetes native platform for training and inferencing across all the DGX machines. The Run.ai acquisition did not change anything, as you leave all the work to the user to integrate with distributed training frameworks like Ray or scalable inference platforms, like KServe/vLLM.