Live data from Hacker News

Making AMD GPUs competitive for LLM inference (2023)

blog.mlc.ai

211–220 of 221 posts

Re: Making AMD GPUs competitive for LLM inference (2023)

#211
post #97

The problem is that performance achievements on AMD consumer-grade GPUs (RX7900XTX) are not representative/transferrable to the Datacenter grade GPUs (MI300X). Consumer GPUs are based on RDNA architecture, while datacenter GPUs are based on the CDNA architecture, and only sometime in ~2026 AMD is expected to release unifying UDNA architecture [1]. At CentML we are currently working on integrating AMD CDNA and HIP sup…

The problem is that the specs of AMD consumer-grade GPUs do not translate to computer performance when you try and chain more than one together. I have 7 NVidia 4090s under my desk happily chugging along on week long training runs. I once managed to get a Radeon VII to run for six hours without shitting itself.

[dead]

Re: Making AMD GPUs competitive for LLM inference (2023)

#212

Earlier quoted context omitted.

AMD decided not to release a high-end GPU this cycle so any investment into 7x00 or 6x00 is going to be wasted as Nvidia 5x00 is likely going to destroy any ROI from the older cards and AMD won't have an answer for at least two years, possibly never due to being non-existing in high-end consumer GPUs usable for compute.

No high-end consumer RDNA4 GPU this cycle. And it's only missing the very high-end model. So we'll still get at least a 7800xt equivalent and whatever CDNA MI models they come out with. The market for the extreme high-end consumer is pretty small, so they're only missing out on clout.

The top-end RDNA4 GPU will have 16GB RAM. That's a massive regression compared to 7900XTX and performance-wise it should be at best at the 7900XTX level. We are discussing AMD cards for LLM inference where VRAM is arguably the most important aspect of a GPU and AMD just threw in the towel for this cycle.

Re: Making AMD GPUs competitive for LLM inference (2023)

#213
post #197

Earlier quoted context omitted.

For context, Alchemist was SIMD8. They made a big deal out of this at the alchemist launch if I recall correctly since they thought it would be more efficient. Unfortunately, it turned out to be less efficient. Tom Petersen did a bunch of interviews right before the Intel B580 launch. In the hardware unboxed interview, he mentioned it, but accidentally misspoke. I must have interpreted his misspeak as meaning games w…

> Having to schedule fewer things is a definite benefit of 32 lanes over a smaller lane count. From a hardware design perspective, it saves you some die size in the scheduler. From a performance perspective, as long as the hardware designer kept 32 in mind, it can schedule 32 lanes and duplicate the signals to the 16 or 8 wide lanes with no loss of performance. > That documentation talks about writing to a temporary…

> From a performance perspective, as long as the hardware designer kept 32 in mind, it can schedule 32 lanes and duplicate the signals to the 16 or 8 wide lanes with no loss of performance.

I was looking at the things that were said for XE2 in Lunar Lake and it appears that the slides suggest that they had special handling to emulate SIMD32 using SIMD16 in hardware, so you might be right.

> So this is a situation where wider lanes actually need more hardware to run at full speed and not having it causes a penalty. I see your point here, but I will note that you can add that criss-cross hardware for 32-wide operations while still having 16-wide be your default.

To go from SIMD8 to SIMD16, Intel halved the number of units while making them double the width. They could have done that again to avoid the need for additional hardware.

I have not seen the Xe2 instruction set to have any hints about how they are doing these operations in their hardware. I am going to leave it at that since I have spent far too much time analyzing the technical marketing for a GPU architecture that I am not likely to use. No matter how well they made it, it just was not scaled up enough to make it interesting to me as a developer that owns a RTX 3090 Ti. I only looked into it as much as I did since I am excited to see Intel moving forward here. That said, if they launched a 48GB variant, I would buy it in a heartbeat and start writing code to run on it.

Re: Making AMD GPUs competitive for LLM inference (2023)

#214
post #127

Earlier quoted context omitted.

> That is like suggesting TSMC has a monopoly. They... do have a monopoly on foundry capacity, especially if you're looking at the most advanced nodes? Nobody's going to Intel or Samsung to build 3nm processors. Hell, there have been whispers over the past month that even Samsung might start outsourcing Exynos to TSMC; Intel already did that with Lunar Lake. Having a monopoly doesn't mean that you are engaging in ant…

This gets at the classic problem in defining a monopoly: how hou define the market. Every company is a monopoly if you define the market narrowly enough. Ford has a monopoly on F150’s. I would argue that defining a semiconductor market in terms of node size is too narrow. Just because TSMC is getting the newest nodes first does not mean they have a monopoly in the semiconductor market. We can play semantics, but for…

> Just because TSMC is getting the newest nodes first does not mean they have a monopoly in the semiconductor market.

Sure. Market research also places them as having somewhere around 65% of worldwide foundry sales [0], with Samsung coming in second place with about 12% (mostly first-party production). Fact is that nobody else comes close to providing real competition for TSMC, so they can charge whatever prices they want, whether you're talking about the 3nm node or the 10nm node.

[0] https://www.counterpointresearch.com/insights/global-semicon...

Rounding out the top five... SMIC (6%) is out of the question unless you're based in China due to various sanctions, UMC (5%) mainly sell decade+-old processes (22nm and larger), and Global Foundries explicitly has abandoned keeping up with the latest technologies.

If you exclude the various Chinese foundries and subtract off Samsung's first-party development, TSMC's share of available foundry capacity for third-party contracts likely grows to 70% or more. At what point do you consider this to be a monopoly? Microsoft Windows has about 72% of desktop OS share.

Re: Making AMD GPUs competitive for LLM inference (2023)

#215

Earlier quoted context omitted.

At a loss seems a bit overly dramatic. I'd guess Nvidia sells SKUs for three times their marginal cost. Intel is probably operating at cost without any hopes of recouping R&D with the current SKUs, but that's reasonable for an aspiring competitor.

It kinda seems they are covering the cost of throwing massive amounts of resources trying to get Arc’s drivers in shape.

The drivers are shared by their iGPUs, so the cost of improving the drivers is likely shared by those.

Re: Making AMD GPUs competitive for LLM inference (2023)

#216

Earlier quoted context omitted.

> Seven and a half years. Seven and a half years was the 2017 Ryzen release date. Zen 1 took them from being completely hopeless to having something competitive but only just, because they were still having the whole thing fabbed by GF. Their revenue didn't exceed what it was in 2011 until 2019 and didn't exceed Intel's until 2022. It's still less than Nvidia, even though AMD is fielding CPUs competitive with Intel a…

It has also been quite a while since Zen+ and Zen 2. Those poured in money, and they absolutely did not need to wait until they had more revenue than some chunk of Intel or until their debt was gone . If you think they got properly started on this in 2022, that's pretty damning. I'm not basing anything on geohotz, just general discussions from people that have tried, and my own experience of trying to get some popula…

> It has also been quite a while since Zen+ and Zen 2. Those poured in money

Zen+ and Zen 2 were released in 2019. Their revenue in 2019 was only 2.5% higher than it was in 2011; adjusted for inflation it was still down more than 10%.

> they absolutely did not need to wait until they had more revenue than some chunk of Intel or until their debt was gone.

The premise of the comparison is that it shows the resources they have available. To make the same level of investment as a bigger company you either have to take it out of profit (not possible when your net profit has a minus sign in front of it or is only a single digit) or you have to make more money first.

And carrying high interest debt when you're now at much lower risk of default is pretty foolish. You'd be paying interest that could be going to R&D. Even if you want to borrow money in order to invest it, the thing to do is to pay back the high interest debt and then borrow the money again now that you can get better terms, which seems to be just what they did.

Re: Making AMD GPUs competitive for LLM inference (2023)

#217

Earlier quoted context omitted.

It has also been quite a while since Zen+ and Zen 2. Those poured in money, and they absolutely did not need to wait until they had more revenue than some chunk of Intel or until their debt was gone . If you think they got properly started on this in 2022, that's pretty damning. I'm not basing anything on geohotz, just general discussions from people that have tried, and my own experience of trying to get some popula…

> It has also been quite a while since Zen+ and Zen 2. Those poured in money Zen+ and Zen 2 were released in 2019. Their revenue in 2019 was only 2.5% higher than it was in 2011; adjusted for inflation it was still down more than 10%. > they absolutely did not need to wait until they had more revenue than some chunk of Intel or until their debt was gone . The premise of the comparison is that it shows the resources t…

I never said anything about wanting them to invest the same amount as nvidia or Intel. I think a handful of extra people could have made a big difference, in particular if some of them had the sole task of bringing their consumer cards into the support list.

It is so bad that they had major cards that were never on the support list for compute.

> You'd be paying interest that could be going to R&D.

Getting people to actually consider your datacenter cards, because they know how to use your cards, will get you more R&D money.

Re: Making AMD GPUs competitive for LLM inference (2023)

#218
post #101

Earlier quoted context omitted.

They've actually been making positive moves with GPUs lately along with a success story for the B580.

B580 being a "success" is purely a business decision as a loss leader to get their name into the market. A larger die on a newer node than either Nvidia or AMD means their per-unit costs are higher, and are selling it at a lower price. That's not a long-term success strategy. Maybe good for getting your name in the conversation, but not sustainable.

Well yes, it's in the name "loss leader". It's not meant to be sustainable. It's meant to get their name out there as a good alternative to Radeon cards for the lower-end GPU market.

Profit can come after positive brand recognition for the product.

Re: Making AMD GPUs competitive for LLM inference (2023)

#219
post #176

I will only consider AMD GPUs for LLM when I can easily make my AMD GPU available within WSL and Docker on Windows. For now, it is as if AMD does not exist in this field for me.

Isn't it already available somehow? I didn't test it seriously, I just needed to quickly run Whisper but $ rocminfo | grep -E "WSL|XTX" WSL environment detected. Marketing Name: AMD Radeon RX 7900 XTX

Interesting. Looking here [1] it seems like this is a thing now. Will dive deeper into this later but looks promising.

[1] https://rocm.docs.amd.com/projects/radeon/en/latest/docs/ins...

Re: Making AMD GPUs competitive for LLM inference (2023)

#220

Earlier quoted context omitted.

The 4090 offers 82.58 teraflops of single-precision performance compared to the Radeon Pro VII's 13.06 teraflops.

On the other hand, for double precision a Radeon Pro VII is many times faster than a RTX 4090 (due to 1:2 vs. 1:64 FP64:FP32 ratio). Moreover, for workloads limited by the memory bandwidth, a Radeon Pro VII and a RTX 4090 will have about the same speed, regardless what kind of computations are performed. It is said that speed limitation by memory bandwidth happens frequently for ML/AI inferencing.

For people wondering:

Titan V: 7.8 TFLOPs

AMD Radeon Pro VII: 6.5 TFLOPs

AMD Radeon VII: 3.52 TFLOPs

4090: 1.3 TFLOPs

Post reply on HN