Live data from Hacker News

Nvidia DGX Spark

nvidia.com

61–70 of 222 posts

Re: Nvidia DGX Spark

#61

I was considering getting an RTX 5090 to run inference on some LLM models, but now I’m wondering if it’s worth paying an extra $2K for this option instead

RTX 5090 for running smaller models.

Then the RTX Pro 6000 for running a little bit larger models (96gb VRAM, but only ~15-20% more perf than 5090).

Some suggest Apple Silicon only for running larger models on a budget because of the unified memory, but the performance won't compare.

Re: Nvidia DGX Spark

#62

FP4-sparse (TFLOPS) | Price | $/TF4s 5090: 3352 | 1999 | 0.60 Thor: 2070 | 3499 | 1.69 Spark: 1000 | 3999 | 4.00 ____________ FP8-dense (TFLOPS) | Price | $/TF8d (4090s have no FP4) 4090 : 661 | 1599 | 2.42 4090 Laptop: 343 | vary | - ____________ Geekbench 6 (compute score) | Price | $/100k 4090: 317800 | 1599 | 503 5090: 387800 | 1999 | 516 M4 Max: 180700 | 1999 | 1106 M3 Ultra: 259700 | 3999 | 1540 ____________ Ap…

where does an RTX Pro 6000 Blackwell fall in this? I feel like that’s the next step up in performance (and about the same price as two Sparks)

Re: Nvidia DGX Spark

#63

FP4-sparse (TFLOPS) | Price | $/TF4s 5090: 3352 | 1999 | 0.60 Thor: 2070 | 3499 | 1.69 Spark: 1000 | 3999 | 4.00 ____________ FP8-dense (TFLOPS) | Price | $/TF8d (4090s have no FP4) 4090 : 661 | 1599 | 2.42 4090 Laptop: 343 | vary | - ____________ Geekbench 6 (compute score) | Price | $/100k 4090: 317800 | 1599 | 503 5090: 387800 | 1999 | 516 M4 Max: 180700 | 1999 | 1106 M3 Ultra: 259700 | 3999 | 1540 ____________ Ap…

How does the process management comparison work for GPU vs full systems?

Re: Nvidia DGX Spark

#64
post #43

Earlier quoted context omitted.

It's the same as Strix Halo and M4 Max that people are going gaga about, so either everyone is wrong or it's fine.

M4 max has more than double the bandwidth. Strix Halo has the same and I agree it’s overrated.

I would expect/hope that DGX would be able to make better use of its bandwidth than the M4 Max. Will need to wait and see benchmarks.

Re: Nvidia DGX Spark

#65

Earlier quoted context omitted.

Again, prompt processing isn't the major problem here. It's bandwidth. 256GB/s bandwidth (maybe ~210 in real world) limits the tokens per second well before prompt processing. Not entirely sure how your ARM statement matters here. This is unified memory.

[flagged]

I like the cut of your jib and your experience matches mine, but without real numbers this is all just piss in the wind (as far as online discussions go).

Re: Nvidia DGX Spark

#66
post #43

The RAM bandwidth is so slow on this that you can barely train or do inference or do anything on it. I think the only use case they have in mind for this is fine tuning pretrained models.

It's the same as Strix Halo and M4 Max that people are going gaga about, so either everyone is wrong or it's fine.

Memory Bandwidth:

Nvidia DGX: 273 GB/s

M4 Max: (up to) 546 GB/s

M3 Ultra: 819 GB/s

RTX 5090: ~1.8 TB/s

RTX PRO 6000 Blackwell: ~1.8 TB/s

Re: Nvidia DGX Spark

#67
post #12

I’m not in this space, so I don’t know what’s normal, but I guess I’m a little surprised to see only 10 gig Ethernet for high speed connectivity. Yeah, it’s miles better than WiFi. But if there was something I’d think maybe benefit from Thunderbolt this would’ve been it. The ability to transfer large models or datasets that way just seems like it would be much faster and a real win for some customers.

This thing has a ConnectX-7, which gives it 2 x 200 Gbps networking. The 10 gig port is far from the fastest network interface on the Spark.

But can you hook that up to a normal PC?

Re: Nvidia DGX Spark

#68

FP4-sparse (TFLOPS) | Price | $/TF4s 5090: 3352 | 1999 | 0.60 Thor: 2070 | 3499 | 1.69 Spark: 1000 | 3999 | 4.00 ____________ FP8-dense (TFLOPS) | Price | $/TF8d (4090s have no FP4) 4090 : 661 | 1599 | 2.42 4090 Laptop: 343 | vary | - ____________ Geekbench 6 (compute score) | Price | $/100k 4090: 317800 | 1599 | 503 5090: 387800 | 1999 | 516 M4 Max: 180700 | 1999 | 1106 M3 Ultra: 259700 | 3999 | 1540 ____________ Ap…

Once the updated Mac Studio with M4/M5 Ultra comes out, pretty much going to make the DGX irrelevant right?

Re: Nvidia DGX Spark

#70
post #56

Most people are missing the point. LLMs are not the be all end all of AI. Even if you were to say memory bandwidth was the problem, there is no consumer grade GPU that can run any SoTA LLM, no matter what you'd have to settle for a more mediocre model. Outside of LLMs, 256 GB/s is not as much of an issue and many people have dealt with less bandwidth for real world use cases.

What other use cases would use 128GB VRAM but not require higher throughput to run at acceptable speeds?
Post reply on HN