Earlier quoted context omitted.
I own one, I don’t feel like the RAM is a huge issue (of course I want 192GB to run something like DS4 Flash). The lack of FP4 and slow memory bandwidth is rough. NVFP4 support is such a huge advantage that I would recommend others to buy a DGX spark over a strix halo if you’re using it purely for AI. Strix halo works better for general computing.
There's a glimmer of hope with ROCmFP4 which seems to double the current throughput: https://github.com/charlie12345/rocmfp4-llama
Hands-On with the AMD Ryzen AI Halo
21–30 of 45 posts
Re: Hands-On with the AMD Ryzen AI Halo
#22Earlier quoted context omitted.
Heads up, you can absolutely run DS4 Flash on a 128gb machine - I have it running on my Strix Halo box right now. https://github.com/antirez/ds4
How has it been and how’s the speed? I read online that the it’s ~200 TPS for PP and ~15 TPS for TG. Unfortunately for those speeds, it’s very very hard to use for agentic stuff.
It definitely seems to be the leader on 'general intelligence' on this hardware from my casual usage, but the newer Qwen or Gemma series models are much more usable speed-wise for agentic use, and often is as good or better than DS4 on that front.
Re: Hands-On with the AMD Ryzen AI Halo
#23Highly recommend lemonade server if you have a strix halo desktop. Been using Qwen3.6-35B @ Q_8 as my main driver and it’s been great with 60 TPS for generation. I occasionally use the 27B @ q6 but only get 20-25 TPS for generation with MTP.
Is there some means to monitor the queries it's sending (or hold and review before transmission) or throttle to avoid triggering abuse thresholds on any single domain?
Re: Hands-On with the AMD Ryzen AI Halo
#24Does anyone else feel like it would be great to be able to purchase $4000 AI box but 128 gigs is not enough. If I spend all that money and it doesn’t really do what I wanted to do, whats the point? It’s kind of like general aviation where you can go buy a Cessna but it’s only going to realistically get you somewhere you could drive anyways but do you really wanna spend that mush cash to get road trip distance at slig…
Yeah, It's be great to have 256GB RAM, but that's really expensive now anyway and there's no way I could get my wife to sign off on spending 5 grand (or more now) on a box with that much memory.
Re: Hands-On with the AMD Ryzen AI Halo
#25Highly recommend lemonade server if you have a strix halo desktop. Been using Qwen3.6-35B @ Q_8 as my main driver and it’s been great with 60 TPS for generation. I occasionally use the 27B @ q6 but only get 20-25 TPS for generation with MTP.
Do you give it access to the internet? Is there some means to monitor the queries it's sending (or hold and review before transmission) or throttle to avoid triggering abuse thresholds on any single domain?
I think it stores basic logs but if you wanted to monitor you’re probably going to want to proxy it with an LLM gateway.
Re: Hands-On with the AMD Ryzen AI Halo
#26Earlier quoted context omitted.
Part of me wonders, would 3d xpoint (if still around) be a viable option? Yeah it is slower than real RAM by a good amount for latency, but you can get similar bandwidth and the cost was history about half of the same size DDR.
I was thought experimenting the other day... ~10 nVME drives striped and running parallel could approximate the memory bandwidth of DDR5 DRAM in a box like this. Like you say, latency wouldn't compete but on raw throughput would be comparable. Not anymore cost effective, I guess, but gets you the ability to work over very large model sizes maybe. But the problem is that tensor matmul etc hardware wouldn't work effect…
Re: Hands-On with the AMD Ryzen AI Halo
#27Until RAM prices drop and can economically get machines with 256GB, 512GB and higher bandwidth... I frankly think the local AI story is going to be still fairly muted for most people. My Spark can do Qwen3.6 MoE A3B at 60 to 70-ish token/second and that's really good, but there's limits the usefulness of that model. It's not useful for coding, in any case. Once people can run something like GLM 5.2 at lower quants (5…
Agree - the 128GB Strix Halo is capable if you use LLMs as assistants , but it's not so good if you use LLMs as agents (or worse, agent teams/swarms) since all of the models that can fit on it are pretty dumb compared to frontier or near-frontier models. You can at best hope for Sonnet-level capabilities. That doesn't mean that local models are useless though! If Mythos/Sol is an ASI that threatens to take your job a…
That memory bandwidth is painful. It's like trying to fill an Olympic swimming pool with a thimble.
It is excellent as a regular PC, though.
Re: Hands-On with the AMD Ryzen AI Halo
#28Does anyone else feel like it would be great to be able to purchase $4000 AI box but 128 gigs is not enough. If I spend all that money and it doesn’t really do what I wanted to do, whats the point? It’s kind of like general aviation where you can go buy a Cessna but it’s only going to realistically get you somewhere you could drive anyways but do you really wanna spend that mush cash to get road trip distance at slig…
Re: Hands-On with the AMD Ryzen AI Halo
#29Re: Hands-On with the AMD Ryzen AI Halo
#30250GB/s on unified memory? That doesn’t sound right, it’s very low