I’m looking forward to trying this currently Strix halo’s npu isn’t accessible if you’re running Linux, and previously I don’t think lemonade was either. If this opens up the npu that would be great! Resolute raccoon is adding npu support as well.
Lemonade by AMD: a fast and open source local LLM server using GPU and NPU
51–60 of 133 posts
Re: Lemonade by AMD: a fast and open source local LLM server using GPU and NPU
#52Note that the NPU models/kernels this uses are proprietary and not available as open source. It would be nice to develop more open support for this hardware.
Re: Lemonade by AMD: a fast and open source local LLM server using GPU and NPU
#53Anyone compare to ollama? I had good success with latest ollama with ROCm 7.4 on 9070 XT a few days ago
better than Vulkan?
Re: Lemonade by AMD: a fast and open source local LLM server using GPU and NPU
#54I have been using lemonade for nearly a year already. On Strix Halo I am using nothing else - although kyuz0's toolboxes are also nice ( https://kyuz0.github.io/amd-strix-halo-toolboxes/ ) Nowadays you get TTS, STT, text & image generation and image editing should also be possible. Besides being able to run via rocm, vulkan or on CPU, GPU and NPU. Quite a lot of options. They have a quite good and pragmatic pace in d…
Have you used it with any agents or claw? If so, which model do you run?
Lemonade has a Web UI to set the context size and llama.cpp args, you need to set context to proper number or just to 0 so that it uses the default. If its too low, it wont work with agentic coding.
I will try some Claw app, but first need to research the field a bit. But I am using different models on Open Web UI. GPT 120B is fast, but also Qwen3.5 27B is fine.
Re: Lemonade by AMD: a fast and open source local LLM server using GPU and NPU
#55This way software adoption will be very limited.
Re: Lemonade by AMD: a fast and open source local LLM server using GPU and NPU
#56Re: Lemonade by AMD: a fast and open source local LLM server using GPU and NPU
#57[flagged]
Re: Lemonade by AMD: a fast and open source local LLM server using GPU and NPU
#58Earlier quoted context omitted.
Have you used it with any agents or claw? If so, which model do you run?
I have two Strix Halo devices at hand. Privately a framework desktop with 128gb and at work 64GB HP notebook. The 64GB machine can load Qwen3.5 30B-A3B, with VSCode it needs a bit of initial prompt processing to initialize all those tools I guess. But the model is fighting with the other resources that I need. So I am not really using it anymore these days, but I want to experiment on my home machine with it. I just…
27B is supposed to be really good but it's so slow I gave up on it (11-12 tg/s at Q4).
Re: Lemonade by AMD: a fast and open source local LLM server using GPU and NPU
#59Is... is this named because they have a lemon they're trying to make the most of?
Re: Lemonade by AMD: a fast and open source local LLM server using GPU and NPU
#60Note that the NPU models/kernels this uses are proprietary and not available as open source. It would be nice to develop more open support for this hardware.
I bought one of their machines to play around with under the expectation that I may never be able to use the NPU for models. But I am still angry to read this anyway.