Live data from Hacker News

Nvidia DGX Spark

nvidia.com

121–130 of 222 posts

Re: Nvidia DGX Spark

#121
post #64

Earlier quoted context omitted.

I would expect/hope that DGX would be able to make better use of its bandwidth than the M4 Max. Will need to wait and see benchmarks.

Matrix vector multiplication for feed forward layers is most of the bandwidth as I understand things, there's not really a way to do it "better", its just a bunch of memory-bound dot products. (Posting this comment in hopes of being corrected and learning something).

The problem is different parts of the SoC (CPU, GPU, NPU) may not actually be able to consume all of the bandwidth available to the system as a whole. This is why you'd need to benchmark - different chips may be able to feed the cores better than others.

Re: Nvidia DGX Spark

#122
post #102

I think it depends on your model size Fits into 32gb: 5090 Fits into 64gb - 96gb: Mac Studio Fits into 128gb: for now 395+ $/token/s, Mac Studio if you don't care about $ but don't have unlimited money for Hxxx This could be great for models that fit 128gb and you want best $/token/s (if it is faster than a 395+).

The 395 although it can be supplied with 128GB can’t use all that for VRAM (unless something has changed in the last couple of weeks).

From YouTube it seems up to 105gb Models disksize work, yes.

Re: Nvidia DGX Spark

#123

FP4-sparse (TFLOPS) | Price | $/TF4s 5090: 3352 | 1999 | 0.60 Thor: 2070 | 3499 | 1.69 Spark: 1000 | 3999 | 4.00 ____________ FP8-dense (TFLOPS) | Price | $/TF8d (4090s have no FP4) 4090 : 661 | 1599 | 2.42 4090 Laptop: 343 | vary | - ____________ Geekbench 6 (compute score) | Price | $/100k 4090: 317800 | 1599 | 503 5090: 387800 | 1999 | 516 M4 Max: 180700 | 1999 | 1106 M3 Ultra: 259700 | 3999 | 1540 ____________ Ap…

As long as you're going to add FP8 dense, you could do the same for the parts mentioned in the FP4 section. Divide by two from dense => sparse, and another two for FP4 => FP8.

That gives you 250 tops of fp8 for Spark.

Re: Nvidia DGX Spark

#124

FP4-sparse (TFLOPS) | Price | $/TF4s 5090: 3352 | 1999 | 0.60 Thor: 2070 | 3499 | 1.69 Spark: 1000 | 3999 | 4.00 ____________ FP8-dense (TFLOPS) | Price | $/TF8d (4090s have no FP4) 4090 : 661 | 1599 | 2.42 4090 Laptop: 343 | vary | - ____________ Geekbench 6 (compute score) | Price | $/100k 4090: 317800 | 1599 | 503 5090: 387800 | 1999 | 516 M4 Max: 180700 | 1999 | 1106 M3 Ultra: 259700 | 3999 | 1540 ____________ Ap…

5090: 32GB RAM (newegg & amazon lowest price seems to be +300 more) 4090: 24GB RAM Thor & Spark: 128GB RAM (probably at least 96GB usable by the GPU if they behave similar to the AMD Strix Halo APU)

True... It would be very interesting to make a comparison of various open models based on token generation speed on these platforms. Presumably starting st some size the larger accessible RAM wins out over raw speed but low VRAM? Although I suppose things like MoE and FP would also matter.

Re: Nvidia DGX Spark

#125
post #110

"developers can prototype, fine-tune, and inference [AI models]"... shouldn't it be infer ?

It should be "run inference on" in my opinion, and would be best shortened IMO to just "prototype, fine-tune, and run".

I'argue that "inference" has taken on a somewhat distinct new meaning in an LLM-context (loosely: running actual tokens through the model) and deviating from the base term to the verb form would make the sentence less clear to me.

Re: Nvidia DGX Spark

#126

Earlier quoted context omitted.

I thought the 6000 was slightly lower throughput than 5090, but obviously has a shitload more RAM.

It's more throughput, but way less value and there's still no NVLink on the 6000. Something like ~4x the price, ~20% more performance, 3x the VRAM. There's two models that go by 6000, the RTX Pro 6000 (Blackwell) is the one that's currently relevant.

the RTX Pro 6000 (Blackwell) does not have NVlink? if so, what the fuck Mr.leather jacket.

Re: Nvidia DGX Spark

#127

FP4-sparse (TFLOPS) | Price | $/TF4s 5090: 3352 | 1999 | 0.60 Thor: 2070 | 3499 | 1.69 Spark: 1000 | 3999 | 4.00 ____________ FP8-dense (TFLOPS) | Price | $/TF8d (4090s have no FP4) 4090 : 661 | 1599 | 2.42 4090 Laptop: 343 | vary | - ____________ Geekbench 6 (compute score) | Price | $/100k 4090: 317800 | 1599 | 503 5090: 387800 | 1999 | 516 M4 Max: 180700 | 1999 | 1106 M3 Ultra: 259700 | 3999 | 1540 ____________ Ap…

Memory is the bottleneck. It limits the size of the models you can run and what you pay for.

  Spark: 128 GB LPDDR5x, unified system memory
  5090 :  32 GB GDDR7,
Model sizes (parameter size)

  Spark: 200B 
  5090 :  12B (raw)

Re: Nvidia DGX Spark

#128

Earlier quoted context omitted.

If that would be true why aren't Mac sales banned in China instead of Nvidia GPUs?

Because Tim bribed Trump with a golden calf, or more seriously it's easier to ban a component and its manufacturer vs broader systems.

Not unprecedented though, Playstation 2 had export restrictions.

But it was a different time. Most policies had some connection to the subject at hand.

Policies today are all about brand Trump and brand MAGA.

Re: Nvidia DGX Spark

#130
post #127

FP4-sparse (TFLOPS) | Price | $/TF4s 5090: 3352 | 1999 | 0.60 Thor: 2070 | 3499 | 1.69 Spark: 1000 | 3999 | 4.00 ____________ FP8-dense (TFLOPS) | Price | $/TF8d (4090s have no FP4) 4090 : 661 | 1599 | 2.42 4090 Laptop: 343 | vary | - ____________ Geekbench 6 (compute score) | Price | $/100k 4090: 317800 | 1599 | 503 5090: 387800 | 1999 | 516 M4 Max: 180700 | 1999 | 1106 M3 Ultra: 259700 | 3999 | 1540 ____________ Ap…

Memory is the bottleneck. It limits the size of the models you can run and what you pay for. Spark: 128 GB LPDDR5x, unified system memory 5090 : 32 GB GDDR7, Model sizes (parameter size) Spark: 200B 5090 : 12B (raw)

[deleted]
Post reply on HN