Live data from Hacker News

Hands-On with the AMD Ryzen AI Halo

microcenter.com

21–30 of 45 posts

Re: Hands-On with the AMD Ryzen AI Halo

#21

Earlier quoted context omitted.

I own one, I don’t feel like the RAM is a huge issue (of course I want 192GB to run something like DS4 Flash). The lack of FP4 and slow memory bandwidth is rough. NVFP4 support is such a huge advantage that I would recommend others to buy a DGX spark over a strix halo if you’re using it purely for AI. Strix halo works better for general computing.

There's a glimmer of hope with ROCmFP4 which seems to double the current throughput: https://github.com/charlie12345/rocmfp4-llama

I saw that but end of the day, the chips themselves don’t have hardware support for FP4. There’s smart ways around this limitation but it will never natively be close to true FP4 performance like MXFP4 and NVFP4 (happy to be proven wrong though).

Re: Hands-On with the AMD Ryzen AI Halo

#22

Earlier quoted context omitted.

Heads up, you can absolutely run DS4 Flash on a 128gb machine - I have it running on my Strix Halo box right now. https://github.com/antirez/ds4

How has it been and how’s the speed? I read online that the it’s ~200 TPS for PP and ~15 TPS for TG. Unfortunately for those speeds, it’s very very hard to use for agentic stuff.

That's pretty accurate to what I've seen, I'd definitely recommend a smaller model for active agentic use on that hardware.

It definitely seems to be the leader on 'general intelligence' on this hardware from my casual usage, but the newer Qwen or Gemma series models are much more usable speed-wise for agentic use, and often is as good or better than DS4 on that front.

Re: Hands-On with the AMD Ryzen AI Halo

#23

Highly recommend lemonade server if you have a strix halo desktop. Been using Qwen3.6-35B @ Q_8 as my main driver and it’s been great with 60 TPS for generation. I occasionally use the 27B @ q6 but only get 20-25 TPS for generation with MTP.

Do you give it access to the internet?

Is there some means to monitor the queries it's sending (or hold and review before transmission) or throttle to avoid triggering abuse thresholds on any single domain?

Re: Hands-On with the AMD Ryzen AI Halo

#24

Does anyone else feel like it would be great to be able to purchase $4000 AI box but 128 gigs is not enough. If I spend all that money and it doesn’t really do what I wanted to do, whats the point? It’s kind of like general aviation where you can go buy a Cessna but it’s only going to realistically get you somewhere you could drive anyways but do you really wanna spend that mush cash to get road trip distance at slig…

I've got a Framework with 128GB. Sure, I'm probably not going to run Deepseek V4 flash (though, I could run the 2-bit quant). But there are a lot of models (especially MOE models) that run fine. Even Qwen3.5-122B in a 4 or 5 bit quant can run. The Qwen3-coder-next (~80B IIRC) runs fine (IIRC I'm running a 6-bit quant of that one). The much vaunted Qwen3.6-27B is runnable, but kind of slow (~20t/s with MTP), though that's not a memory limitation issue.

Yeah, It's be great to have 256GB RAM, but that's really expensive now anyway and there's no way I could get my wife to sign off on spending 5 grand (or more now) on a box with that much memory.

Re: Hands-On with the AMD Ryzen AI Halo

#25

Highly recommend lemonade server if you have a strix halo desktop. Been using Qwen3.6-35B @ Q_8 as my main driver and it’s been great with 60 TPS for generation. I occasionally use the 27B @ q6 but only get 20-25 TPS for generation with MTP.

Do you give it access to the internet? Is there some means to monitor the queries it's sending (or hold and review before transmission) or throttle to avoid triggering abuse thresholds on any single domain?

Lemonade proxies your request to llama.cpp that it installs and manages. Is out of the box all local LLMs, but you can connect it to an API endpoint (supports OpenAI compatible or Anthropic compatible).

I think it stores basic logs but if you wanted to monitor you’re probably going to want to proxy it with an LLM gateway.

Re: Hands-On with the AMD Ryzen AI Halo

#26

Earlier quoted context omitted.

Part of me wonders, would 3d xpoint (if still around) be a viable option? Yeah it is slower than real RAM by a good amount for latency, but you can get similar bandwidth and the cost was history about half of the same size DDR.

I was thought experimenting the other day... ~10 nVME drives striped and running parallel could approximate the memory bandwidth of DDR5 DRAM in a box like this. Like you say, latency wouldn't compete but on raw throughput would be comparable. Not anymore cost effective, I guess, but gets you the ability to work over very large model sizes maybe. But the problem is that tensor matmul etc hardware wouldn't work effect…

I'll just leave this here: "Achieving 11M IOPS and 66 GB/S IO on a Single ThreadRipper Workstation" (2021), https://news.ycombinator.com/item?id=25956670 / https://tanelpoder.com/posts/11m-iops-with-10-ssds-on-amd-th...

Re: Hands-On with the AMD Ryzen AI Halo

#27

Until RAM prices drop and can economically get machines with 256GB, 512GB and higher bandwidth... I frankly think the local AI story is going to be still fairly muted for most people. My Spark can do Qwen3.6 MoE A3B at 60 to 70-ish token/second and that's really good, but there's limits the usefulness of that model. It's not useful for coding, in any case. Once people can run something like GLM 5.2 at lower quants (5…

Agree - the 128GB Strix Halo is capable if you use LLMs as assistants , but it's not so good if you use LLMs as agents (or worse, agent teams/swarms) since all of the models that can fit on it are pretty dumb compared to frontier or near-frontier models. You can at best hope for Sonnet-level capabilities. That doesn't mean that local models are useless though! If Mythos/Sol is an ASI that threatens to take your job a…

Exactly this. I own the Framework desktop board. I knew all of its limitations before I bought it, and it's ok to play with on a hobby level, but it isn't much more than a Radio Shack toy.

That memory bandwidth is painful. It's like trying to fill an Olympic swimming pool with a thimble.

It is excellent as a regular PC, though.

Re: Hands-On with the AMD Ryzen AI Halo

#28

Does anyone else feel like it would be great to be able to purchase $4000 AI box but 128 gigs is not enough. If I spend all that money and it doesn’t really do what I wanted to do, whats the point? It’s kind of like general aviation where you can go buy a Cessna but it’s only going to realistically get you somewhere you could drive anyways but do you really wanna spend that mush cash to get road trip distance at slig…

Agreed, the Strix Halo doesn't feel good enough for $4k. It was supposed to be $2k (and was so until 6-12 months ago), which felt like a great deal. Not the best AI chip but you get what you pay for. A tinkerer's dream that could maybe even fit into a birthday gift budget for a lucky teenager. I hate to say it, but I hope they fail with their $4k box.
Post reply on HN