Live data from Hacker News

Nvidia DGX Spark and Apple Mac Studio = 4x Faster LLM Inference with EXO 1.0

blog.exolabs.net

11–20 of 22 posts

Re: Nvidia DGX Spark and Apple Mac Studio = 4x Faster LLM Inference with EXO 1.0

#12
post #10
post #3

Are you using USB-C for networking between the Spark and the Mac?

IP over thunderbolt is definitely a thing, don't know whether IP over USB is also a thing. USB4x2 or TB5 can do 80Gib/s symmetrical or 120+40 asymmetrical (and boy is this a poster child for the asymmetrical setup). The Mac definitely supports that fine, so, as long as the Spark plays nice, USB is actually a legitimately decent choice.

USB4 was based on Thunderbolt3

Yes, it's a thing that works.

Re: Nvidia DGX Spark and Apple Mac Studio = 4x Faster LLM Inference with EXO 1.0

#15
post #13
post #7

This is really cool! Now I'm trying to stop myself from finding an excuse to spend upwards of $30k on compute hardware...

if you have $30k to spare, I'm sure there are better options

Yeah, a couple of RTX Pro 6000 cards would blow this away and still leave him with money to spare.

Re: Nvidia DGX Spark and Apple Mac Studio = 4x Faster LLM Inference with EXO 1.0

#17

It’s really sad that exo went private.

How do you know this happened? I thought it was an abandoned project until I saw this post. I've been diligently checking weekly for new releases but nothing for almost a year...

Re: Nvidia DGX Spark and Apple Mac Studio = 4x Faster LLM Inference with EXO 1.0

#19
post #14

Wouldn't this restrict memory to 128GB, wasting M3 Ultra potential?

Blog author here. Actually, no. The model can be streamed into the DGX Spark, so we can run prefill of models much larger than 128GB (e.g. DeepSeek R1) on the DGX Spark. This feature is coming to EXO 1.0 which will be open-sourced soonTM.

Re: Nvidia DGX Spark and Apple Mac Studio = 4x Faster LLM Inference with EXO 1.0

#20
post #2

Very cool, using the DGX like an “AI eGPU.” I wonder if this could also benefit stuff like Stable Diffusion/WAN etc?

Yes, these models are mostly compute-bound so benefit even more from the compute on the DGX Spark.
Post reply on HN