Live data from Hacker News

Making AMD GPUs competitive for LLM inference (2023)

blog.mlc.ai

101–110 of 221 posts

Re: Making AMD GPUs competitive for LLM inference (2023)

#101

Earlier quoted context omitted.

More cynical take: this would be a bad strategy, because Intel hasn't shown much competence in its leadership for a long time, especially in regards to GPUs.

They've actually been making positive moves with GPUs lately along with a success story for the B580.

B580 being a "success" is purely a business decision as a loss leader to get their name into the market. A larger die on a newer node than either Nvidia or AMD means their per-unit costs are higher, and are selling it at a lower price.

That's not a long-term success strategy. Maybe good for getting your name in the conversation, but not sustainable.

Re: Making AMD GPUs competitive for LLM inference (2023)

#102

Great, I have yet to understand why does not the ML community really push or move away from CUDA? To me, it feel like a dinosaur move to build on top of CUDA which is screaming proprietary nothing about it is open source or cross platform. The reason why I say its dinosaur is, imagine, we as a dev community continued to build on top of Flash or Microsoft Silverlight... LLM and ML has been out for quiet a while, with…

For me personally, hacking together projects as a hobbiest, 2 reasons : 1. It just works. When i tried to build things on Intel Arcs, i spent way more hours bikeshedding ipex and driver issues than developing 2. LLMs seem to have more cuda code in their training data. I can leverage claude and 4o to help me build things with cuda, but trying to get them to help me do the same things on ipex just doesn't work. I'd ver…

What kind of model learn and what's its token output on intel gpu's?

Re: Making AMD GPUs competitive for LLM inference (2023)

#103
post #97

The problem is that performance achievements on AMD consumer-grade GPUs (RX7900XTX) are not representative/transferrable to the Datacenter grade GPUs (MI300X). Consumer GPUs are based on RDNA architecture, while datacenter GPUs are based on the CDNA architecture, and only sometime in ~2026 AMD is expected to release unifying UDNA architecture [1]. At CentML we are currently working on integrating AMD CDNA and HIP sup…

The problem is that the specs of AMD consumer-grade GPUs do not translate to computer performance when you try and chain more than one together. I have 7 NVidia 4090s under my desk happily chugging along on week long training runs. I once managed to get a Radeon VII to run for six hours without shitting itself.

Wow, are these 7 RTX 4090s in a single setup? Care to share more how you build it (case, cooling, power, ..)?

Re: Making AMD GPUs competitive for LLM inference (2023)

#104

Earlier quoted context omitted.

Except I never hear complaints about CUDA from a quality perspective. The complaints are always about lock in to the best GPUs on the market. The desire to shift away is to make cheaper hardware with inferior software quality more usable. Flash was an abomination, CUDA is not.

Flash was popular because it was an attractive platform for the developer. Back then there was no HTML5 and browsers didn't otherwise support a lot of the things Flash did. Flash Player was an abomination, it was crashy and full of security vulnerabilities, but that was a problem for the user rather than the developer and it was the developer choosing what to use to make the site. This is pretty much exactly what hap…

> users have to use expensive hardware with proprietary drivers/firmware

What do you mean by that? People trying to run their own models are not “the users” they are a tiny insignificant niche segment.

Re: Making AMD GPUs competitive for LLM inference (2023)

#105

The problem is that performance achievements on AMD consumer-grade GPUs (RX7900XTX) are not representative/transferrable to the Datacenter grade GPUs (MI300X). Consumer GPUs are based on RDNA architecture, while datacenter GPUs are based on the CDNA architecture, and only sometime in ~2026 AMD is expected to release unifying UDNA architecture [1]. At CentML we are currently working on integrating AMD CDNA and HIP sup…

It looks like AMD's CDNA gpu's are supported by Mesa, which ought to suffice for Vulkan Compute and SYCL support. So there should be ways to run ML workloads on the hardware without going through HIP/ROCm.

Re: Making AMD GPUs competitive for LLM inference (2023)

#106

Earlier quoted context omitted.

And I'm supposed to believe that HN is this amazing platform for technology and science discussions, totally unlike its peers...

Maybe be the change you want to see and tell us what the real story is?

We seem to disagree on what the change in the world I'd like to see is like, which is a real shocker I'm sure.

Personally, I think that's when somebody who has no real information to contribute doesn't try to pretend that they do.

So thanks for the offer, but I think I'm already delivering on that realm.

Re: Making AMD GPUs competitive for LLM inference (2023)

#107
post #56

Earlier quoted context omitted.

And I'm supposed to believe that HN is this amazing platform for technology and science discussions, totally unlike its peers...

I don't really care what you believe. Everyone whose dug deep into what AMD is doing has left in disgust if they are lucky and bankruptcy if they are not. If I can save someone else from wasting $100,000 on hardware and six months of their life then my post has done more good than the AMD marketing department ever will.

I'd be very concerned if somebody makes a $100K decision based on a comment where the author couldn't even differentiate between the words "constitutionally" and "institutionally", while providing as much substance as any other random techbro on any random forum and being overwhelmingly oblivious to it.

Re: Making AMD GPUs competitive for LLM inference (2023)

#108

Earlier quoted context omitted.

Gigabytes per second? What is this, bandwidth for ants? My years old pleb tier non-HBM GPU has more than 4 times the bandwidth you would get from a PCIe Gen 7 x16 link, which doesn't even officially exist yet.

> 4 times the bandwidth you would get from a PCIe Gen 7 x16 link So you have a full terabyte per second of bandwidth? What GPU is that? (The 64GB/s number is an x4 link. If you meant you have over four times that, then it sounds like CXL would be pretty competitive.)

RTX 4090 comes to mind. Dunno that I'd consider that a "years old pleb tier non-HBM GPU" though.

Re: Making AMD GPUs competitive for LLM inference (2023)

#109
post #86

Earlier quoted context omitted.

Yeah but MLID says they are losing money on every one and have been winding down the internal development resources. That doesn't bode well for the future. I want to believe he's wrong, but on the parts of his show where I am in a position to verify, he generally checks out. Whatever the opposite of Gell-Mann Amnesia is, he's got it going for him.

The die size of the B580 is 272 mm2, which is a lot of silicon for $249. The performance of the GPU is good for its price but bad for its die size. Manufacturing cost is closely tied to die size. 272 mm2 puts the B580 in the same league as the Radeon 7700XT, a $449 card, and the GeForce 4070 Super, which is $599. The idea that Intel is selling these cards at a loss sounds reasonable to me.

Though you assume the prices of the competition are reasonable. There are plenty of reasons for them not to be. Availability issues, lack of competition, other more lucrative avenues etc.

Intel has neither, or at least not as much of them.

Re: Making AMD GPUs competitive for LLM inference (2023)

#110
post #52
post #16

Earlier quoted context omitted.

Tinygrad was another one, but they ended up getting frustrated with AMD and semi-pivoted to Nvidia.

This is discussed in the lex Friedman episode. AMD’s own demo would kernel panic when run in a loop [1]. [1] https://youtube.com/watch?v=dNrTrx42DGQ&t=3218

Interesting. I wonder if focusing on GPUs and CPUs is something that requires two companies instead of one, whether the concentration of resources just leads to one arm of your company being much better than the other.
Post reply on HN