Live data from Hacker News

Making AMD GPUs competitive for LLM inference (2023)

blog.mlc.ai

91–100 of 221 posts

Re: Making AMD GPUs competitive for LLM inference (2023)

#93
Just an FYI, this is writeup from August 2023 and a lot has changed (for the better!) for RDNA3 AI/ML support.

That being said, I did some very recent inference testing on an W7900 (using the same testing methodology used by Embedded LLM's recent post to compare to vLLM's recently added Radeon GGUF support [1]) and MLC continues to perform quite well. On Llama 3.1 8B, MLC's q4f16_1 (4.21MB weights) performed +35% faster than llama.cpp w/ Q4_K_M w/ their ROCm/HIP backend (4.30MB weights, 2% size difference).

That makes MLC still the generally fastest standalone inference engine for RDNA3 by a country mile. However, you have much less flexibility with quants and by and large have to compile your own for every model, so llama.cpp is probably still more flexible for general use. Also llama.cpp's (recently added to llama-server) speculative decoding can also give some pretty sizable performance gains. Using a 70B Q4_K_M + 1B Q8_0 draft model improves output token throughput by 59% on the same ShareGPT testing. I've also been running tests with Qwen2.5-Coder and using a 0.5-3B draft model for speculative decoding gives even bigger gains on average (depends highly on acceptance rate).

Note, I think for local use, vLLM GGUF is still not suitable at all. When testing w/ a 70B Q4_K_M model (only 40GB), loading, engine warmup, and graph compilation took on avg 40 minutes. llama.cpp takes 7-8s to load the same model.

At this point for RDNA3, basically everything I need works/runs for my use cases (primarily LLM development and local inferencing), but almost always slower than an RTX 3090/A6000 Ampere (a new 24GB 7900 XTX is $850 atm, used or refurbished 24 GB RTX 3090s are in in the same ballpark, about $800 atm; a new 48GB W7900 goes for $3600 while an 48GB A6000 (Ampere) goes for $4600). The efficiency gains can be sizable. Eg, on my standard llama-bench test w/ llama2-7b-q4_0, the RTX 3090 gets a tg128 of 168 t/s while the 7900 XTX only gets 118 t/s even though both have similar memory bandwidth (936.2 GB/s vs 960 GB/s). It's also worth noting that since the beginning of the year, the llama.cpp CUDA implementation has gotten almost 25% faster, while the ROCm version's performance has stayed static.

There is an actively (solo dev) maintained fork of llama.cpp that sticks close to HEAD but basically applies a rocWMMA patch that can improve performance if you use the llama.cpp FA (still performs worse than w/ FA disabled) and in certain long-context inference generations (on llama-bench and w/ this ShareGPT serving test you won't see much difference) here: https://github.com/hjc4869/llama.cpp - The fact that no one from AMD has shown any interest in helping improve llama.cpp performance (despite often citing llama.cpp-based apps in marketing/blog posts, etc is disappointing ... but sadly on brand for AMD GPUs).

Anyway, for those interested in more information and testing for AI/ML setup for RDNA3 (and AMD ROCm in general), I keep a doc with lots of details here: https://llm-tracker.info/howto/AMD-GPUs

[1] https://embeddedllm.com/blog/vllm-now-supports-running-gguf-...

Re: Making AMD GPUs competitive for LLM inference (2023)

#94
post #47

Earlier quoted context omitted.

Cynical take: Try to get acquired by Intel for Arc.

Intel is in a vastly better shape than AMD, they have the software pretty much nailed down.

Someone never used intel killer wifi software.

Re: Making AMD GPUs competitive for LLM inference (2023)

#95
post #47

Earlier quoted context omitted.

Cynical take: Try to get acquired by Intel for Arc.

Intel is in a vastly better shape than AMD, they have the software pretty much nailed down.

I've recently been poking around with Intel oneAPI and IPEX-LLM. While there are things that I find refreshing (like their ability to actually respond to bug reports in a timely manner, or at all) on a whole, support/maturity actually doesn't match the current state of ROCm.

PyTorch requires it's own support kit separate from the oneAPI Toolkit (and runs slightly different versions of everything), the vLLM xpu support doesn't work - both source and the docker failed to build/run for me. The IPEX-LLM whisper support is completely borked, etc, etc.

Re: Making AMD GPUs competitive for LLM inference (2023)

#96
post #56

Earlier quoted context omitted.

I don't really care what you believe. Everyone whose dug deep into what AMD is doing has left in disgust if they are lucky and bankruptcy if they are not. If I can save someone else from wasting $100,000 on hardware and six months of their life then my post has done more good than the AMD marketing department ever will.

> If I can save someone else from wasting $100,000 on hardware and six months of their life then my post has done more good than the AMD marketing department ever will. This seems like unuseful advice if you've already given up on them. You tried it and at some point in the past it wasn't ready. But by not being ready they're losing money, so they have a direct incentive to fix it. Which would take a certain amount o…

I've tried it three times.

I've seen people try it every six months for two decades now.

At some point you just have to accept that AMD is not a serious company, but is a second rate copycat and there is no way to change that without firing everyone from middle management up.

I'm deeply worried about stagnation in the CPU space now that they are top dog and Intel is dead in the water.

Here's hoping China and Risk V save us.

>Meanwhile the people who attempt it apparently seem to get acquired by Nvidia

Everyone I've seen base jumping has gotten a sponsorship from redbull, ergo. everyone should basejump.

Ignore the red smears around the parking lot.

Re: Making AMD GPUs competitive for LLM inference (2023)

#97

The problem is that performance achievements on AMD consumer-grade GPUs (RX7900XTX) are not representative/transferrable to the Datacenter grade GPUs (MI300X). Consumer GPUs are based on RDNA architecture, while datacenter GPUs are based on the CDNA architecture, and only sometime in ~2026 AMD is expected to release unifying UDNA architecture [1]. At CentML we are currently working on integrating AMD CDNA and HIP sup…

The problem is that the specs of AMD consumer-grade GPUs do not translate to computer performance when you try and chain more than one together.

I have 7 NVidia 4090s under my desk happily chugging along on week long training runs. I once managed to get a Radeon VII to run for six hours without shitting itself.

Re: Making AMD GPUs competitive for LLM inference (2023)

#98

Earlier quoted context omitted.

Yeah but MLID says they are losing money on every one and have been winding down the internal development resources. That doesn't bode well for the future. I want to believe he's wrong, but on the parts of his show where I am in a position to verify, he generally checks out. Whatever the opposite of Gell-Mann Amnesia is, he's got it going for him.

Wait, are they losing money on every one in the sense that they haven't broken even on research and development yet? Or in the sense that they cost more to manufacture than they're sold at? Because one is much worse than the other.

They're trying to unseat Radeon as the budget card. That means making a more enticing offer than AMD for a temporary period of time.

Re: Making AMD GPUs competitive for LLM inference (2023)

#100

Earlier quoted context omitted.

Is there no hope for AMD anymore? After George Hotz/Tinygrad gave up on AMD I feel there’s no realistic chance of using their chips to break the CUDA dominance.

The world is bigger than AMD and Nvidia. Plenty of interesting new AI-tuned non-GPU accelerators coming online.

I hope, name some NPU who can run a 70B model..
Post reply on HN