Live data from Hacker News

Nvidia DGX Spark as a daily driver

daniel.lawrence.lu

31–40 of 85 posts

Re: Nvidia DGX Spark as a daily driver

#31
post #26

I have been running uConsoles with CM5 (2712 and 3588 with 16GB RAM) for 6 months as daily drivers. They are ~$500* and present the same ARM problems/opportunities. But they are completely silent (no fan, the case is the heat sink). My 6600(3050) desktop from 2016(2024) with replaced SSD(2021)/RAM(2025) (they age like milk) now gets little use and M$ will soon sleep with the fishes. *Hard to get now as the 3588 that…

Got 2 of these early on. They are wonderfully designed, unfortunately I've been caught in the "building things" trap for few months now and they have been relegated to being used as retro gaming/ computing learning boxes for my sons. They really don't appreciate it at all (yet)....but I can hope they will remember it in the future. Teaching them to get to terminal and run the emulator was great fun...reminded me of MSDoS and the hours of troubleshooting to run games with limited memory and drivers back in the day. I just worry that with LLMs the whole point of teaching them basic terminal/troubleshooting skills might be lost soon. We will see.

Re: Nvidia DGX Spark as a daily driver

#32

This is an interesting review. I have a Chinese strix halo box that's isnt available in the west (favm faex1) I've been able to do some ok graphical gen, or some decent agentic tasks as a fallback for when some of the APIs are overloaded during business hours, but nothing amazing for sure, and also not both at the same time. But here's the thing...it cost me 1800usd two months ago....and it's runs x86. I am strugglin…

It's simple, the 395+ Max Strix Halo you bought for $1800 is now a ~$4000 build (at least the AMD AI dev unit). If only we had time travel right? Either way, the Nvidia unit comes with Connect-X 7. That may or may not matter to you, but the hardware for that isn't cheap. In general the Nvidia cards also have better support for models. I know AMD is trying to catch up but anything except their datacenter cards do not seem to be getting a lot of attention.

Re: Nvidia DGX Spark as a daily driver

#33
Please don’t buy a DGX Spark unless all three of these are true:

  - You value simplicity more than performance or price-to-performance.
  - You accept that the hardware will depreciate rapidly.
  - You’re prepared to buy two or four of them.
OR:

  - You want to run frontier models right now as cheaply as possible
  - You want to run high-parameter models on a 15a breaker/line
Otherwise, get a normal, high-bandwidth GPU.

A single Spark gives you roughly 115 GB of usable memory compared with the 24–32 GB found on many lower-cost GPUs. It's certainly a big increase, but in practice it does not unlock dramatically better models.

  - One Spark: More memory, but mostly enough for poor-quality, extremely low-bit quants of larger models.  
  - Two Sparks: Enough for mid-tier parameter models at reasonable quants, such as DeepSeek V4 Flash and HY3.
  - Four Sparks: Enough for GLM 5.2 at a reasonable quant.  You'll need a $1000+ switch too.
The problem is that Sparks are slow compared with almost everything else in their price range. Many factors affect inference speed, but memory bandwidth is one of the biggest. A $4,000-plus DGX Spark provides only 273 GB/s.

  | GPU                 | Memory bandwidth |           VRAM | Approx. price  |
  | ------------------- | ---------------: | -------------: | ------------:  |
  | DGX Spark           |         273 GB/s | ~115 GB usable |        $4,000+ |
  | RTX 5060            |         448 GB/s |          16 GB |          $600  |
  | Radeon AI Pro R9700 |         640 GB/s |          32 GB |        $1,200  |
  | RTX 4000 Pro        |         672 GB/s |          24 GB |        $2,300  |
  | RTX 4500 Pro        |         896 GB/s |          32 GB |        $3,500  |
  | RTX 3090            |         936 GB/s |          24 GB |        $1,200  |
  | RTX 5000 Pro        |       1,344 GB/s |          48 GB |        $6,000  |
  | RTX 5090            |       1,792 GB/s |          32 GB |        $4,000  |
  | RTX 6000 Pro        |       1,792 GB/s |          96 GB |       $12,000  |
Yes, the Spark has substantially more memory. But going from roughly 24 GB to 115 GB does not necessarily unlock substantially better model quality. In many cases, it only lets you load heavily compressed 2-bit versions of larger models, such as DeepSeek V4 Flash, with serious quality degradation.

24–32 GB is currently a sweet spot. Models such as Qwen 3.6 27B and 35B-A3B:

  - Perform far above what their parameter counts suggest.
  - Fit comfortably within 24–32 GB of VRAM at reasonable quantization levels.
A 4-bit quant of Qwen 3.6 27b (18 GB) will out-perform a 2-bit quant of DeepSeek v4 Flash (90gb).

Instead of the Spark, if I had a roughly $4,000 budget...

Assuming I already had a reasonably modern desktop:

  - One RTX 5090, RTX 5000 Pro, or RTX 4500 Pro.
  - Two RTX 3090s, RTX 4000 Pros, or R9700s, provided the motherboard can bifurcate two physical x16 slots into x8/x8.
If I were building a system from scratch:

  - A DDR4- or PCIe 4.0-era consumer CPU and motherboard that supports x8/x8 bifurcation.
  - Two RTX 3090s, RTX 4000 Pros, or R9700s.
If I were already planning to buy a new Mac:

  - A MacBook Pro M5 with 64 GB or 128 GB of unified memory.
For context, these are the systems I currently run:

  - EPYC Turin with four RTX 6000 Pro Max-Qs.
  - EPYC Milan with four RTX 3090s.
  - AM4 with two RTX 3090s.
  - AM4 with two RTX 3090s.
  - Intel Raptor Lake with two RTX 5060 Ti.
  - MacBook Pro M3 128GB Unified

Re: Nvidia DGX Spark as a daily driver

#34

This is an interesting review. I have a Chinese strix halo box that's isnt available in the west (favm faex1) I've been able to do some ok graphical gen, or some decent agentic tasks as a fallback for when some of the APIs are overloaded during business hours, but nothing amazing for sure, and also not both at the same time. But here's the thing...it cost me 1800usd two months ago....and it's runs x86. I am strugglin…

> isnt available in the west [...] Can anyone explain the allure of the Nvidia box The first part might answer the second one. Otherwise, the lack of CUDA and the nvidia ecosystem of tooling could also explain why it doesn't seem so interesting for AI tasks.

But strix halo boxes themselves are available, just not that one. And despite my concerns about what's said about Cuda and ROCm I have never had a problem running any model, for image or text or voice, the community has done great work in making things work. So the point of the question stands.

It's also interesting that the most high end Chinese equipment, both prosumer things like these boxes but also the Huawei professional stack is just not available in the places it would be most appreciated. Not sure if thats china tit for tat, or western "we don't want your commie hardware anyways"

But for a lot of people it sucks cause nobody should be paying 4.2k for this product. The value isn't there.

Re: Nvidia DGX Spark as a daily driver

#35

This is an interesting review. I have a Chinese strix halo box that's isnt available in the west (favm faex1) I've been able to do some ok graphical gen, or some decent agentic tasks as a fallback for when some of the APIs are overloaded during business hours, but nothing amazing for sure, and also not both at the same time. But here's the thing...it cost me 1800usd two months ago....and it's runs x86. I am strugglin…

> isnt available in the west [...] Can anyone explain the allure of the Nvidia box The first part might answer the second one. Otherwise, the lack of CUDA and the nvidia ecosystem of tooling could also explain why it doesn't seem so interesting for AI tasks.

Yep, stuff is more likely to "just work" on NVidia. Example: pytorch

In most benchmarks, the Spark is also faster at the prompt processing / prefill phase.

Re: Nvidia DGX Spark as a daily driver

#36
post #9

I'm genuinely disappointed with my Spark. I don't know how anyone can claim it performs decently with LLMs or diffusion models. Back when I worked in VFX in the early 2000s, we had a saying: "Render time is coffee time" and if you try to run this thing with a usable context size, you'll be drinking a lot of coffee. Most of the optimizations it relies on for inference simply aren't available for training, so it crawls…

I have access to both a RTX 5090 PC with 64gb of ram and a Spark 128gb, the performance of the Spark has been highly disappointing.

I prefer 1000x the RTX one, even with 64gb of ram.

Re: Nvidia DGX Spark as a daily driver

#37

This is an interesting review. I have a Chinese strix halo box that's isnt available in the west (favm faex1) I've been able to do some ok graphical gen, or some decent agentic tasks as a fallback for when some of the APIs are overloaded during business hours, but nothing amazing for sure, and also not both at the same time. But here's the thing...it cost me 1800usd two months ago....and it's runs x86. I am strugglin…

It's simple, the 395+ Max Strix Halo you bought for $1800 is now a ~$4000 build (at least the AMD AI dev unit). If only we had time travel right? Either way, the Nvidia unit comes with Connect-X 7. That may or may not matter to you, but the hardware for that isn't cheap. In general the Nvidia cards also have better support for models. I know AMD is trying to catch up but anything except their datacenter cards do not…

AMD is barely trying. The 395+ Max Strix Halo was launched January 2025.

gfx1151 was not listed in the ROCm compatibility matrix for ROCm 7.2.4 [0]. This is the previous version of ROCm.

It's only finally received support in ROCm 7.14.0 [1]! It literally just started receiving support last week.

[0] https://rocm.docs.amd.com/en/docs-7.2.4/compatibility/compat...

[1] https://rocm.docs.amd.com/en/docs-7.14.0/compatibility/compa...

Re: Nvidia DGX Spark as a daily driver

#38
post #22
post #17

Neat blog! I was intrigued by this bullet point mentioned in passing: >my four hard drive USB 3.2 ZFS raidz2 array with four 24 TB drives Can you speak more about this? Which USB array did you choose? How well does it work? I've been slowly planning a transition away from my power-hungry surplus enterprise gear in the 19" rack towards a smaller, quieter, lower power setup ... but storage is the real kicker right now.…

It's an Orico 9948C3 with four Seagate Barracuda 24TB drives. They were on sale last year [1]. Unfortunately, the enclosure doesn't work super well on Linux. There is a weird bug where the drives don't enumerate when I boot up my computer. This happens on both my x86_64 AMD machine running Linux, and on the DGX Spark. The solution is... simply power cycle the enclosure a couple of times by toggling the power button o…

Yeah... that's been my past experience with USB docks, at least any with more than one slot. I've never had great luck with them. "Can recover from power outage without being touched" is a key requirement for my NAS so I'll give that one a pass and stick with my "SATA drives directly attached to a SAS controller" strategy for now. Thanks for the reply.

As for the 24/7 use, yeah so be it. The I in RAID stands for Inexpensive. If they fail after 10 years at 24/7, so be it. I have drive level redundancy and frequent offsite backups of anything critical.

Re: Nvidia DGX Spark as a daily driver

#40

Earlier quoted context omitted.

> isnt available in the west [...] Can anyone explain the allure of the Nvidia box The first part might answer the second one. Otherwise, the lack of CUDA and the nvidia ecosystem of tooling could also explain why it doesn't seem so interesting for AI tasks.

But strix halo boxes themselves are available, just not that one. And despite my concerns about what's said about Cuda and ROCm I have never had a problem running any model, for image or text or voice, the community has done great work in making things work. So the point of the question stands. It's also interesting that the most high end Chinese equipment, both prosumer things like these boxes but also the Huawei pr…

> I have never had a problem running any model, for image or text or voice, the community has done great work in making things work

There is a whole world of other tooling and stuff that isn't just for hobbyists to run inference with ML models, but also how to do profiling, debugging and gathering data when you run distributed workloads, and so on. The nsight toolkit seems miles ahead of the competition on other platforms, as just one example.

Post reply on HN