Live data from Hacker News

Nvidia DGX Spark

nvidia.com

41–50 of 222 posts

Re: Nvidia DGX Spark

#42

I was considering getting an RTX 5090 to run inference on some LLM models, but now I’m wondering if it’s worth paying an extra $2K for this option instead

RTX 5090 is about as good as it gets for home use. Its inference speeds are extremely fast.

The limiting factor is going to be the VRAM on the 5090, but nvidia intentionally makes trying to break the 32GB barrier extremely painful - they want companies to buy their $20,000 GPUs to run inference for larger models.

Re: Nvidia DGX Spark

#43

The RAM bandwidth is so slow on this that you can barely train or do inference or do anything on it. I think the only use case they have in mind for this is fine tuning pretrained models.

It's the same as Strix Halo and M4 Max that people are going gaga about, so either everyone is wrong or it's fine.

Re: Nvidia DGX Spark

#45

Earlier quoted context omitted.

Their prompt processing speeds are absolutely abysmal They are not. This is Blackwell with Tensor cores. Bandwidth is the problem here.

They're abysmal compared to anything dedicated at any reasonable batch size because of both bandwidth and compute, not sure why you're wording this like it disagrees with what I said. I've run inference workloads on a GH200 which is an entire H100 attached to an ARM processor and the moment offloading is involved speeds tank to Mac Mini-like speeds, which is similarly mostly a toy when it comes to AI.

Again, prompt processing isn't the major problem here. It's bandwidth. 256GB/s bandwidth (maybe ~210 in real world) limits the tokens per second well before prompt processing.

Not entirely sure how your ARM statement matters here. This is unified memory.

Re: Nvidia DGX Spark

#48
post #43

The RAM bandwidth is so slow on this that you can barely train or do inference or do anything on it. I think the only use case they have in mind for this is fine tuning pretrained models.

It's the same as Strix Halo and M4 Max that people are going gaga about, so either everyone is wrong or it's fine.

The other ones are not framed as an “AI Supercomputer on your desk”, but instead are framed as powerful computers that can also handle AI workloads.

Re: Nvidia DGX Spark

#49
FP4-sparse (TFLOPS) | Price | $/TF4s

5090: 3352 | 1999 | 0.60

Thor: 2070 | 3499 | 1.69

Spark: 1000 | 3999 | 4.00

____________

FP8-dense (TFLOPS) | Price | $/TF8d (4090s have no FP4)

4090 : 661 | 1599 | 2.42

4090 Laptop: 343 | vary | -

____________

Geekbench 6 (compute score) | Price | $/100k

4090: 317800 | 1599 | 503

5090: 387800 | 1999 | 516

M4 Max: 180700 | 1999 | 1106

M3 Ultra: 259700 | 3999 | 1540

____________

Apple NPU TOPS (not GPU-comparable)

M4 Max: 38

M3 Ultra: 36

Re: Nvidia DGX Spark

#50

FP4-sparse (TFLOPS) | Price | $/TF4s 5090: 3352 | 1999 | 0.60 Thor: 2070 | 3499 | 1.69 Spark: 1000 | 3999 | 4.00 ____________ FP8-dense (TFLOPS) | Price | $/TF8d (4090s have no FP4) 4090 : 661 | 1599 | 2.42 4090 Laptop: 343 | vary | - ____________ Geekbench 6 (compute score) | Price | $/100k 4090: 317800 | 1599 | 503 5090: 387800 | 1999 | 516 M4 Max: 180700 | 1999 | 1106 M3 Ultra: 259700 | 3999 | 1540 ____________ Ap…

It's not good value when you put it like that. It doesn't have a lot of compute and bandwidth. What it has is the ability to run DGX software for CUDA devs I guess. Not a great inference machine either.
Post reply on HN