Live data from Hacker News

Making AMD GPUs competitive for LLM inference (2023)

blog.mlc.ai

71–80 of 221 posts

Re: Making AMD GPUs competitive for LLM inference (2023)

#71
post #47

Earlier quoted context omitted.

Cynical take: Try to get acquired by Intel for Arc.

More cynical take: this would be a bad strategy, because Intel hasn't shown much competence in its leadership for a long time, especially in regards to GPUs.

They've actually been making positive moves with GPUs lately along with a success story for the B580.

Re: Making AMD GPUs competitive for LLM inference (2023)

#72
The problem is that performance achievements on AMD consumer-grade GPUs (RX7900XTX) are not representative/transferrable to the Datacenter grade GPUs (MI300X). Consumer GPUs are based on RDNA architecture, while datacenter GPUs are based on the CDNA architecture, and only sometime in ~2026 AMD is expected to release unifying UDNA architecture [1]. At CentML we are currently working on integrating AMD CDNA and HIP support into our Hidet deep learning compiler [2], which will also power inference workloads for all Nvidia GPUs, AMD GPUs, Google TPU and AWS Inf2 chips on our platform [3]

[1] https://www.jonpeddie.com/news/amd-to-integrate-cdna-and-rdn.... [2] https://centml.ai/hidet/ [3] https://centml.ai/platform/

Re: Making AMD GPUs competitive for LLM inference (2023)

#73
post #65

Earlier quoted context omitted.

Except I never hear complaints about CUDA from a quality perspective. The complaints are always about lock in to the best GPUs on the market. The desire to shift away is to make cheaper hardware with inferior software quality more usable. Flash was an abomination, CUDA is not.

Maybe the situation has gotten better in recent years, but my experience with Nvidia toolchains was a complete nightmare back in 2018.

The cuda situation is definitely better. The nvidia struggles are now with the higher-level software they’re pushing (triton, tensor-llm, riva, etc), tools that are the most performant option when they work, but a garbage developer experience when you step outside the golden path

Re: Making AMD GPUs competitive for LLM inference (2023)

#74
post #9

I have come across quite few startups who are trying a similar idea: break the nvidia monopoly by utilizing AMD GPUs (for inference at least): Felafax, Lamini, tensorwave (partially), SlashML. Even saw optimistic claims like CUDA moat is only 18 months deep from some of them [1]. Let's see. [1] https://www.linkedin.com/feed/update/urn:li:activity:7275885...

From Lamini, we have a private AMD GPU cluster, ready to serve any one who want to try MI300x or MI250 with inference and tuning.

We just onboarded a customer to move from openai API to on-prem solution, currently evaluating MI300x for inference.

Email me at my profile email.

Re: Making AMD GPUs competitive for LLM inference (2023)

#75

Earlier quoted context omitted.

More cynical take: this would be a bad strategy, because Intel hasn't shown much competence in its leadership for a long time, especially in regards to GPUs.

They've actually been making positive moves with GPUs lately along with a success story for the B580.

Yeah but MLID says they are losing money on every one and have been winding down the internal development resources. That doesn't bode well for the future.

I want to believe he's wrong, but on the parts of his show where I am in a position to verify, he generally checks out. Whatever the opposite of Gell-Mann Amnesia is, he's got it going for him.

Re: Making AMD GPUs competitive for LLM inference (2023)

#76
post #58
post #55

Earlier quoted context omitted.

Everything is comparative. AMD isn't perfect. As an Ex Shareholder I have argued they did well partly because of Intel's downfall. In terms of execution it is far from perfect. But Nvidia is a different beast. It is a bit like Apple in the late 00s where you take business, forecast, marketing, operation, software, hardware, sales etc You take any part of it and they are all industry leading. And having industry leadi…

Yeah, no. We are in the middle of a monopoly squeeze by NVidia on the most innovative part of the economy right now. I expect the DOJ to hit them harder than they did MS in the 90s given the bullshit they are pulling and the drag on the economy they are causing. By comparison if AMD could write a driver that didn't shit itself when it had to multiply more than two matrices in a row they'd be selling cards faster than…

>We are in the middle of a monopoly squeeze by NVidia on the most innovative part of the economy right now.

I am not sure which part of Nvidia is monopoly. That is like suggesting TSMC has a monopoly.

Re: Making AMD GPUs competitive for LLM inference (2023)

#78

Earlier quoted context omitted.

They've actually been making positive moves with GPUs lately along with a success story for the B580.

Yeah but MLID says they are losing money on every one and have been winding down the internal development resources. That doesn't bode well for the future. I want to believe he's wrong, but on the parts of his show where I am in a position to verify, he generally checks out. Whatever the opposite of Gell-Mann Amnesia is, he's got it going for him.

Wait, are they losing money on every one in the sense that they haven't broken even on research and development yet? Or in the sense that they cost more to manufacture than they're sold at? Because one is much worse than the other.

Re: Making AMD GPUs competitive for LLM inference (2023)

#79

Earlier quoted context omitted.

Compute Express Link (CXL) should mostly solve limited RAM with CPU: 1) Compute Express Link (CXL): https://en.wikipedia.org/wiki/Compute_Express_Link PCIe vs. CXL for Memory and Storage: https://news.ycombinator.com/item?id=38125885

Gigabytes per second? What is this, bandwidth for ants? My years old pleb tier non-HBM GPU has more than 4 times the bandwidth you would get from a PCIe Gen 7 x16 link, which doesn't even officially exist yet.

> 4 times the bandwidth you would get from a PCIe Gen 7 x16 link

So you have a full terabyte per second of bandwidth? What GPU is that?

(The 64GB/s number is an x4 link. If you meant you have over four times that, then it sounds like CXL would be pretty competitive.)

Re: Making AMD GPUs competitive for LLM inference (2023)

#80

Earlier quoted context omitted.

They've actually been making positive moves with GPUs lately along with a success story for the B580.

Yeah but MLID says they are losing money on every one and have been winding down the internal development resources. That doesn't bode well for the future. I want to believe he's wrong, but on the parts of his show where I am in a position to verify, he generally checks out. Whatever the opposite of Gell-Mann Amnesia is, he's got it going for him.

MLID on Intel is starting to become the same as UserBenchmark on AMD (except for the generally reputable sources)... he's beginning to sound like he simply wants Intel to fail, to my insider-info-lacking ears. For competition's sake I really hope that MLID has it wrong (at least the opining about the imminent failure of Intel's GPU division), and that the B series will encourage Intel to push farther to spark more competition in the GPU space.
Post reply on HN