Live data from Hacker News

Nvidia DGX Spark

nvidia.com

141–150 of 222 posts

Re: Nvidia DGX Spark

#141

Earlier quoted context omitted.

I mean the spark is $3,999 and current M3 Max 28-Core CPU 60-Core GPU is the same price. I would expect the refreshed studio will stay around the same price.

In Germany the 96gb version is 5000 EUR and the 256gb version is 7000 EUR (no 128gb available as far as I can see).

Are you comparing prices with or without taxes? US usually prices without and EU with.

Re: Nvidia DGX Spark

#142

Earlier quoted context omitted.

If that would be true why aren't Mac sales banned in China instead of Nvidia GPUs?

Because Tim bribed Trump with a golden calf, or more seriously it's easier to ban a component and its manufacturer vs broader systems.

[flagged]

Re: Nvidia DGX Spark

#143
post #24

The mainstream options seem to be Ryzen AI Max 395+, ~120 tops (fp8?), 128GB RAM, $1999 Nvidia DGX Spark, ~1000 tops fp4, 128GB RAM, $3999 Mac Studio max spec, ~120 tflops (fp16?), 512GB RAM, 3x bandwidth, $9499 DGX Spark appears to potentially offer the most token per second, but less useful/value as everyday pc.

> Ryzen AI Max 395+, ~120 tops (fp8?), 128GB RAM, $1999 Just got my Framework PC last week. It's easy to setup to run LLMs locally - you have to use Fedora 42, though, because it has the latest drivers. It was super easy to get qwen3-coder-30b (8 bit quant) running in LMStudio at 36 tok/sec.

I'm pretty new to this, so if I wanted to benchmark my current hardware and compare to your results what would be the best way to do that?

I'm looking at going for a Framework Desktop and would like to know what kind of performance gain I'd get over the current hardware I have, which so far I have a "feel" for the performance of from running Ollama and OpenWebUI, but no hard numbers.

Re: Nvidia DGX Spark

#145
According to Wendell from Level1Techs, the now-launched Jetson Thor uses a Linux Kernel built by Nvidia, on Ubuntu 20.04 [0]. So I assume getting upgrades will have the same feel as those Chinese SBC's like from Radxa or cheap Android devices.

I wonder if this also applies to this DGX Spark. I hope not.

[0] https://www.youtube.com/watch?v=cgnKUUcCKcs&t=669s

Re: Nvidia DGX Spark

#146
post #127

FP4-sparse (TFLOPS) | Price | $/TF4s 5090: 3352 | 1999 | 0.60 Thor: 2070 | 3499 | 1.69 Spark: 1000 | 3999 | 4.00 ____________ FP8-dense (TFLOPS) | Price | $/TF8d (4090s have no FP4) 4090 : 661 | 1599 | 2.42 4090 Laptop: 343 | vary | - ____________ Geekbench 6 (compute score) | Price | $/100k 4090: 317800 | 1599 | 503 5090: 387800 | 1999 | 516 M4 Max: 180700 | 1999 | 1106 M3 Ultra: 259700 | 3999 | 1540 ____________ Ap…

Memory is the bottleneck. It limits the size of the models you can run and what you pay for. Spark: 128 GB LPDDR5x, unified system memory 5090 : 32 GB GDDR7, Model sizes (parameter size) Spark: 200B 5090 : 12B (raw)

That's very true and what's segmenting the market, but I don't understand why you're saying the 5090 supports only 12B model when it can go up to 50-60B (= a bit less than 64B to leave room for inference) as it supports FP4 as well.

Re: Nvidia DGX Spark

#148
post #127

Earlier quoted context omitted.

Memory is the bottleneck. It limits the size of the models you can run and what you pay for. Spark: 128 GB LPDDR5x, unified system memory 5090 : 32 GB GDDR7, Model sizes (parameter size) Spark: 200B 5090 : 12B (raw)

That's very true and what's segmenting the market, but I don't understand why you're saying the 5090 supports only 12B model when it can go up to 50-60B (= a bit less than 64B to leave room for inference) as it supports FP4 as well.

Its for comparison using raw, non optimized models. Both can do much better when you optimize for inference.

Information is in the ratio of these numbers. They stay the same.

Re: Nvidia DGX Spark

#149
post #148

Earlier quoted context omitted.

That's very true and what's segmenting the market, but I don't understand why you're saying the 5090 supports only 12B model when it can go up to 50-60B (= a bit less than 64B to leave room for inference) as it supports FP4 as well.

Its for comparison using raw, non optimized models. Both can do much better when you optimize for inference. Information is in the ratio of these numbers. They stay the same.

Ok then just to clarify: you can fit 4x larger models on the Spark vs 5090, not 17x.

Re: Nvidia DGX Spark

#150

FP4-sparse (TFLOPS) | Price | $/TF4s 5090: 3352 | 1999 | 0.60 Thor: 2070 | 3499 | 1.69 Spark: 1000 | 3999 | 4.00 ____________ FP8-dense (TFLOPS) | Price | $/TF8d (4090s have no FP4) 4090 : 661 | 1599 | 2.42 4090 Laptop: 343 | vary | - ____________ Geekbench 6 (compute score) | Price | $/100k 4090: 317800 | 1599 | 503 5090: 387800 | 1999 | 516 M4 Max: 180700 | 1999 | 1106 M3 Ultra: 259700 | 3999 | 1540 ____________ Ap…

It's not good value when you put it like that. It doesn't have a lot of compute and bandwidth. What it has is the ability to run DGX software for CUDA devs I guess. Not a great inference machine either.

It's great at one thing: memory. And that's interesting because memory is a commodity, but they still make bank on just being able to access it.
Post reply on HN