Live data from Hacker News

Making AMD GPUs competitive for LLM inference (2023)

blog.mlc.ai

201–210 of 221 posts

Re: Making AMD GPUs competitive for LLM inference (2023)

#201
post #127
post #76

Earlier quoted context omitted.

>We are in the middle of a monopoly squeeze by NVidia on the most innovative part of the economy right now. I am not sure which part of Nvidia is monopoly. That is like suggesting TSMC has a monopoly.

> That is like suggesting TSMC has a monopoly. They... do have a monopoly on foundry capacity, especially if you're looking at the most advanced nodes? Nobody's going to Intel or Samsung to build 3nm processors. Hell, there have been whispers over the past month that even Samsung might start outsourcing Exynos to TSMC; Intel already did that with Lunar Lake. Having a monopoly doesn't mean that you are engaging in ant…

This gets at the classic problem in defining a monopoly: how hou define the market. Every company is a monopoly if you define the market narrowly enough. Ford has a monopoly on F150’s.

I would argue that defining a semiconductor market in terms of node size is too narrow. Just because TSMC is getting the newest nodes first does not mean they have a monopoly in the semiconductor market. We can play semantics, but for any meaningful discussion of monopolistic behaviors, a temporary technical advantage seems a poor way to define the term.

Re: Making AMD GPUs competitive for LLM inference (2023)

#202
post #97

The problem is that performance achievements on AMD consumer-grade GPUs (RX7900XTX) are not representative/transferrable to the Datacenter grade GPUs (MI300X). Consumer GPUs are based on RDNA architecture, while datacenter GPUs are based on the CDNA architecture, and only sometime in ~2026 AMD is expected to release unifying UDNA architecture [1]. At CentML we are currently working on integrating AMD CDNA and HIP sup…

The problem is that the specs of AMD consumer-grade GPUs do not translate to computer performance when you try and chain more than one together. I have 7 NVidia 4090s under my desk happily chugging along on week long training runs. I once managed to get a Radeon VII to run for six hours without shitting itself.

What software stack you use for training?

Re: Making AMD GPUs competitive for LLM inference (2023)

#203

Earlier quoted context omitted.

AMD GPUs are becoming a serious contender for LLM inference. vLLM is already showing impressive performance on AMD [1], even with consumer-grade Radeon cards (even support GGUF) [2]. This could be a game-changer for folks who want to run LLMs without shelling out for expensive NVIDIA hardware. [1] https://blog.vllm.ai/2024/10/23/vllm-serving-amd.html [2] https://embeddedllm.com/blog/vllm-now-supports-running-gguf-...

AMD decided not to release a high-end GPU this cycle so any investment into 7x00 or 6x00 is going to be wasted as Nvidia 5x00 is likely going to destroy any ROI from the older cards and AMD won't have an answer for at least two years, possibly never due to being non-existing in high-end consumer GPUs usable for compute.

No high-end consumer RDNA4 GPU this cycle. And it's only missing the very high-end model. So we'll still get at least a 7800xt equivalent and whatever CDNA MI models they come out with.

The market for the extreme high-end consumer is pretty small, so they're only missing out on clout.

Re: Making AMD GPUs competitive for LLM inference (2023)

#204

Earlier quoted context omitted.

> At some point you just have to accept that AMD is not a serious company, but is a second rate copycat and there is no way to change that without firing everyone from middle management up. AMD has always punched above their weight. Historically their problem was that they were the much smaller company and under heavy resource constraints. Around the turn of the century the Athlon was faster than the Pentium III and…

> So until quite recently the answer to the question "why didn't they do X?" was obvious. They didn't have the money. But now they do. Seven and a half years. The excuse is threadbare at best. They are not doing a reasonable job of making compute work off the shelf.

> Seven and a half years.

Seven and a half years was the 2017 Ryzen release date. Zen 1 took them from being completely hopeless to having something competitive but only just, because they were still having the whole thing fabbed by GF. Their revenue didn't exceed what it was in 2011 until 2019 and didn't exceed Intel's until 2022. It's still less than Nvidia, even though AMD is fielding CPUs competitive with Intel and GPUs competitive with Nvidia at the same time.

They had a pretty good revenue jump in 2021 but much of that was used to pay down debt, because debt taken on when you're almost bankrupt tends to have unfavorable terms. So it wasn't until somewhere in 2022 that they finally got free of GF and the old debt and could start doing something about this. But then it takes some amount of time to actually do it, and you would expect to be seeing the results of that approximately right now. Which seems like a silly time to stop looking.

Also, somewhat counterintuitively, George Hotz et al seem to be employing a strategy in the nature of "say bad things about them in public to shame them into improving", which has the dual result of actually working (they fix a lot of the things he's complaining about) but also making people think that things are worse than they are because there is now a large public archive of rants about things they've already fixed. It's not clear if this is the company not providing a good mechanism for people to complain about things like that in private and have them fixed promptly so it doesn't have take media attention to make it happen, or it's George Hotz seeking publicity as is his custom, or some combination of both.

Re: Making AMD GPUs competitive for LLM inference (2023)

#205
post #9

I have come across quite few startups who are trying a similar idea: break the nvidia monopoly by utilizing AMD GPUs (for inference at least): Felafax, Lamini, tensorwave (partially), SlashML. Even saw optimistic claims like CUDA moat is only 18 months deep from some of them [1]. Let's see. [1] https://www.linkedin.com/feed/update/urn:li:activity:7275885...

AMD GPUs are becoming a serious contender for LLM inference. vLLM is already showing impressive performance on AMD [1], even with consumer-grade Radeon cards (even support GGUF) [2]. This could be a game-changer for folks who want to run LLMs without shelling out for expensive NVIDIA hardware. [1] https://blog.vllm.ai/2024/10/23/vllm-serving-amd.html [2] https://embeddedllm.com/blog/vllm-now-supports-running-gguf-...

These blog posts were written based on my company, Hot Aisle, donating the compute. =) Super proud of being able to support this.

Re: Making AMD GPUs competitive for LLM inference (2023)

#206
post #9

I have come across quite few startups who are trying a similar idea: break the nvidia monopoly by utilizing AMD GPUs (for inference at least): Felafax, Lamini, tensorwave (partially), SlashML. Even saw optimistic claims like CUDA moat is only 18 months deep from some of them [1]. Let's see. [1] https://www.linkedin.com/feed/update/urn:li:activity:7275885...

Hot Aisle (my company) has MI300x compute available for rent too! =)

Re: Making AMD GPUs competitive for LLM inference (2023)

#207
post #97

The problem is that performance achievements on AMD consumer-grade GPUs (RX7900XTX) are not representative/transferrable to the Datacenter grade GPUs (MI300X). Consumer GPUs are based on RDNA architecture, while datacenter GPUs are based on the CDNA architecture, and only sometime in ~2026 AMD is expected to release unifying UDNA architecture [1]. At CentML we are currently working on integrating AMD CDNA and HIP sup…

The problem is that the specs of AMD consumer-grade GPUs do not translate to computer performance when you try and chain more than one together. I have 7 NVidia 4090s under my desk happily chugging along on week long training runs. I once managed to get a Radeon VII to run for six hours without shitting itself.

How do you manage heat? I'm looking at a hashcat build with a few 5090, and water cooling seems to be the sensible solution if we scale beyond two cards.

Re: Making AMD GPUs competitive for LLM inference (2023)

#208
post #197

Earlier quoted context omitted.

> Tom Petersen made a big deal about 16-lane SIMD in Battlemage [...] Where? The only mention I see in that interview is him briefly saying they have native 16 with "simple emulation" for 32 because some games want 32. I see no mention of or comparison to 8. And it doesn't make sense to me that switching to actual 32 would be an improvement. Wider means less flexible here. I'd say a more accurate framing is whether t…

For context, Alchemist was SIMD8. They made a big deal out of this at the alchemist launch if I recall correctly since they thought it would be more efficient. Unfortunately, it turned out to be less efficient. Tom Petersen did a bunch of interviews right before the Intel B580 launch. In the hardware unboxed interview, he mentioned it, but accidentally misspoke. I must have interpreted his misspeak as meaning games w…

There is a typo in the Tom Petersen quote. He said “compute shader”, not “computer shader”. Autocorrect changed it when I had transcribed it and I did not catch this during the edit window.

Re: Making AMD GPUs competitive for LLM inference (2023)

#209
post #197

Earlier quoted context omitted.

> Tom Petersen made a big deal about 16-lane SIMD in Battlemage [...] Where? The only mention I see in that interview is him briefly saying they have native 16 with "simple emulation" for 32 because some games want 32. I see no mention of or comparison to 8. And it doesn't make sense to me that switching to actual 32 would be an improvement. Wider means less flexible here. I'd say a more accurate framing is whether t…

For context, Alchemist was SIMD8. They made a big deal out of this at the alchemist launch if I recall correctly since they thought it would be more efficient. Unfortunately, it turned out to be less efficient. Tom Petersen did a bunch of interviews right before the Intel B580 launch. In the hardware unboxed interview, he mentioned it, but accidentally misspoke. I must have interpreted his misspeak as meaning games w…

> Having to schedule fewer things is a definite benefit of 32 lanes over a smaller lane count.

From a hardware design perspective, it saves you some die size in the scheduler.

From a performance perspective, as long as the hardware designer kept 32 in mind, it can schedule 32 lanes and duplicate the signals to the 16 or 8 wide lanes with no loss of performance.

> That documentation talks about writing to a temporary location and reading form a temporary location in order to do cross lane operations.

> If games’ shaders are written with an assumption that SIMD32 is used, then native SIMD32 is going to be more performant than native SIMD16 because of faster cross lane operations.

So this is a situation where wider lanes actually need more hardware to run at full speed and not having it causes a penalty. I see your point here, but I will note that you can add that criss-cross hardware for 32-wide operations while still having 16-wide be your default.

Re: Making AMD GPUs competitive for LLM inference (2023)

#210

Earlier quoted context omitted.

> So until quite recently the answer to the question "why didn't they do X?" was obvious. They didn't have the money. But now they do. Seven and a half years. The excuse is threadbare at best. They are not doing a reasonable job of making compute work off the shelf.

> Seven and a half years. Seven and a half years was the 2017 Ryzen release date. Zen 1 took them from being completely hopeless to having something competitive but only just, because they were still having the whole thing fabbed by GF. Their revenue didn't exceed what it was in 2011 until 2019 and didn't exceed Intel's until 2022. It's still less than Nvidia, even though AMD is fielding CPUs competitive with Intel a…

It has also been quite a while since Zen+ and Zen 2. Those poured in money, and they absolutely did not need to wait until they had more revenue than some chunk of Intel or until their debt was gone. If you think they got properly started on this in 2022, that's pretty damning.

I'm not basing anything on geohotz, just general discussions from people that have tried, and my own experience of trying to get some popular compute code bases to run. It has been so lacking compared to AMD's own support for games. I'm not going to be "silly" and "stop looking" going forward, but I'm not going to forget how long my card was largely abandoned. It went directly from "not ready yet, working on it" to "obsolete, maybe dregs will be added later".

Post reply on HN