Live data from Hacker News

Lemonade by AMD: a fast and open source local LLM server using GPU and NPU

lemonade-server.ai

61–70 of 133 posts

Re: Lemonade by AMD: a fast and open source local LLM server using GPU and NPU

#61

Earlier quoted context omitted.

better than Vulkan?

In my experience using llama.cpp (which ollama uses internally) on a Strix Halo, whether ROCm or Vulkan performs better really depends on the model and it's usually within 10%. I have access to an RX 7900 XT I should compare to though.

Perhaps I should just google it, but I'm under the impression that ollama uses llama.cpp internally, not the other way around.

Thanks for that data point I should experiment with ROCm

Re: Lemonade by AMD: a fast and open source local LLM server using GPU and NPU

#62
post #2

Anyone compare to ollama? I had good success with latest ollama with ROCm 7.4 on 9070 XT a few days ago

I just compared this on my Mac book M1 Max 64GB RAM with the following:

Model: qwen3.59b Prompt: "Hey, tell me a story about going to space"

Ollama completed in about 1:44 Lemonade completed in about 1:14

So it seems faster in this very limited test.

Re: Lemonade by AMD: a fast and open source local LLM server using GPU and NPU

#64

Is... is this named because they have a lemon they're trying to make the most of?

Lemonsqueeze was considered too violent

If you run it in a cluster, does it become a Lemon Party?

Re: Lemonade by AMD: a fast and open source local LLM server using GPU and NPU

#65

Earlier quoted context omitted.

In my experience using llama.cpp (which ollama uses internally) on a Strix Halo, whether ROCm or Vulkan performs better really depends on the model and it's usually within 10%. I have access to an RX 7900 XT I should compare to though.

Perhaps I should just google it, but I'm under the impression that ollama uses llama.cpp internally, not the other way around. Thanks for that data point I should experiment with ROCm

I meant ollama uses llama.cpp internally. Sorry for the confusion.

Re: Lemonade by AMD: a fast and open source local LLM server using GPU and NPU

#67

Just in case anyone isn't aware. NPUs are low power, slow, and meant for small models.

I wonder what was the imagined use case? TBH I was seriously thinking about buying a framework desktop but the NPU put me off.. I don't get why I should have to pay money for a bunch of silicon that doesn't do anything. And now that there's some software support... it still doesn't do anything? Why does it even exist at all then?

Re: Lemonade by AMD: a fast and open source local LLM server using GPU and NPU

#68

Earlier quoted context omitted.

better than Vulkan?

[flagged]

I was talking about ROCm vs Vulkan. On AMD GPUs, Vulkan has been commonly recognized as the faster API for some time. Both have been slower than CUDA due to most of the hosting projects focusing entirely on Nvidia. Parent post seemed to indicate that newer ROCm releases are better.

Re: Lemonade by AMD: a fast and open source local LLM server using GPU and NPU

#70

Earlier quoted context omitted.

I think saying "L-L-M" sounds kind of like "lemon," so this is an LLM-aid (sounds like lemonade).

so obvious and yet I didn't connect the dots. thank you

wait until you discover the LuLuleMonade -connection /s
Post reply on HN