Hands-On with the AMD Ryzen AI Halo
microcenter.com
Hands-On with the AMD Ryzen AI Halo
1–10 of 45 posts
Re: Hands-On with the AMD Ryzen AI Halo
#2My Spark can do Qwen3.6 MoE A3B at 60 to 70-ish token/second and that's really good, but there's limits the usefulness of that model. It's not useful for coding, in any case.
Once people can run something like GLM 5.2 at lower quants (512GB could do a passable job), then I think the story changes.
Whether we ever see DRAM as cheap as it was ever again, I don't know.
Re: Hands-On with the AMD Ryzen AI Halo
#3Re: Hands-On with the AMD Ryzen AI Halo
#4Until RAM prices drop and can economically get machines with 256GB, 512GB and higher bandwidth... I frankly think the local AI story is going to be still fairly muted for most people. My Spark can do Qwen3.6 MoE A3B at 60 to 70-ish token/second and that's really good, but there's limits the usefulness of that model. It's not useful for coding, in any case. Once people can run something like GLM 5.2 at lower quants (5…
That doesn't mean that local models are useless though! If Mythos/Sol is an ASI that threatens to take your job and turn you into paperclips, then Qwen/Gemma is an old-fashioned office secretary that loyally helps you with tasks but doesn't have a good grasp of details. Every white-collar worker 50 years ago would have killed to have a hard-working personal secretary.
Re: Hands-On with the AMD Ryzen AI Halo
#5Until RAM prices drop and can economically get machines with 256GB, 512GB and higher bandwidth... I frankly think the local AI story is going to be still fairly muted for most people. My Spark can do Qwen3.6 MoE A3B at 60 to 70-ish token/second and that's really good, but there's limits the usefulness of that model. It's not useful for coding, in any case. Once people can run something like GLM 5.2 at lower quants (5…
Yeah it is slower than real RAM by a good amount for latency, but you can get similar bandwidth and the cost was history about half of the same size DDR.
Re: Hands-On with the AMD Ryzen AI Halo
#6Re: Hands-On with the AMD Ryzen AI Halo
#7Until RAM prices drop and can economically get machines with 256GB, 512GB and higher bandwidth... I frankly think the local AI story is going to be still fairly muted for most people. My Spark can do Qwen3.6 MoE A3B at 60 to 70-ish token/second and that's really good, but there's limits the usefulness of that model. It's not useful for coding, in any case. Once people can run something like GLM 5.2 at lower quants (5…
Part of me wonders, would 3d xpoint (if still around) be a viable option? Yeah it is slower than real RAM by a good amount for latency, but you can get similar bandwidth and the cost was history about half of the same size DDR.
Re: Hands-On with the AMD Ryzen AI Halo
#8Until RAM prices drop and can economically get machines with 256GB, 512GB and higher bandwidth... I frankly think the local AI story is going to be still fairly muted for most people. My Spark can do Qwen3.6 MoE A3B at 60 to 70-ish token/second and that's really good, but there's limits the usefulness of that model. It's not useful for coding, in any case. Once people can run something like GLM 5.2 at lower quants (5…
Part of me wonders, would 3d xpoint (if still around) be a viable option? Yeah it is slower than real RAM by a good amount for latency, but you can get similar bandwidth and the cost was history about half of the same size DDR.
Not anymore cost effective, I guess, but gets you the ability to work over very large model sizes maybe. But the problem is that tensor matmul etc hardware wouldn't work effectively with it.
Useful for KVCache though.