Nvidia DGX Spark
91–100 of 222 posts
Re: Nvidia DGX Spark
#92The mainstream options seem to be Ryzen AI Max 395+, ~120 tops (fp8?), 128GB RAM, $1999 Nvidia DGX Spark, ~1000 tops fp4, 128GB RAM, $3999 Mac Studio max spec, ~120 tflops (fp16?), 512GB RAM, 3x bandwidth, $9499 DGX Spark appears to potentially offer the most token per second, but less useful/value as everyday pc.
Re: Nvidia DGX Spark
#93Paper launch. The people I know there who I have asked about it haven't even seen one yet
Re: Nvidia DGX Spark
#94Earlier quoted context omitted.
This thing has a ConnectX-7, which gives it 2 x 200 Gbps networking. The 10 gig port is far from the fastest network interface on the Spark.
But can you hook that up to a normal PC?
Re: Nvidia DGX Spark
#95Re: Nvidia DGX Spark
#96Paper launch. The people I know there who I have asked about it haven't even seen one yet
Ordered one in spring. Delivery time was pushed from July to September. Apparently they had a bug in the HDMI output.
Re: Nvidia DGX Spark
#97The mainstream options seem to be Ryzen AI Max 395+, ~120 tops (fp8?), 128GB RAM, $1999 Nvidia DGX Spark, ~1000 tops fp4, 128GB RAM, $3999 Mac Studio max spec, ~120 tflops (fp16?), 512GB RAM, 3x bandwidth, $9499 DGX Spark appears to potentially offer the most token per second, but less useful/value as everyday pc.
> Ryzen AI Max 395+, ~120 tops (fp8?), 128GB RAM, $1999 Just got my Framework PC last week. It's easy to setup to run LLMs locally - you have to use Fedora 42, though, because it has the latest drivers. It was super easy to get qwen3-coder-30b (8 bit quant) running in LMStudio at 36 tok/sec.
Re: Nvidia DGX Spark
#98Re: Nvidia DGX Spark
#99Earlier quoted context omitted.
M4 max has more than double the bandwidth. Strix Halo has the same and I agree it’s overrated.
I would expect/hope that DGX would be able to make better use of its bandwidth than the M4 Max. Will need to wait and see benchmarks.
Re: Nvidia DGX Spark
#100Earlier quoted context omitted.
Again, prompt processing isn't the major problem here. It's bandwidth. 256GB/s bandwidth (maybe ~210 in real world) limits the tokens per second well before prompt processing. Not entirely sure how your ARM statement matters here. This is unified memory.
[flagged]
I suspect that you’re running a very large model like DeepSeek in coherent memory?
Keep in mind that this little DGX only has 128GB which means it can run fairly small models such as qwen3 coder where prompt processing is not an issue.
I’m not doubting your experience with GH200 but it doesn’t seem relevant here because the bandwidth for Spark is the bottleneck well before the prompt processing.