Live data from Hacker News

Making AMD GPUs competitive for LLM inference (2023)

blog.mlc.ai

81–90 of 221 posts

Re: Making AMD GPUs competitive for LLM inference (2023)

#81
post #37

Earlier quoted context omitted.

I went through the thread. There’s an argument to be made in firing Su for being so spaced out as to miss an op for their own CUDA for free.

Not remotely, how did you get to that idea?

Kids this days (shakes fist)

tl;dr there's a non-unsubstantial # of people who learn a lot from geohot. I'd say about 3% of people here will be confused if you thought of him as less than a top technical expert across many comp sci fields.

And he did the geohot thing recently, way tl;dr: acted like there was a scandal being covered up by AMD around drivers that was causing them to "lose" to nVidia.

He then framed AMD not engaging with him on this topic as further covering-up and choosing to lose.

So if you're of a certain set of experiences, you see an anodyne quote from the CEO that would have been utterly unsurprising dating back to when ATI was still a company, and you'd read it as the CEO breezily admitting in public that geohot was right about how there was malfeasance, followed by a cover up, implying extreme dereliction of duty, because she either helped or didn't realize till now.

I'd argue this is partially due to stonk-ification of discussions, there was a vague, yet often communicated, sense there was something illegal happening. Idea was it was financial dereliction of duty to shareholders.

Re: Making AMD GPUs competitive for LLM inference (2023)

#82
post #47

Earlier quoted context omitted.

Could be trying to make themselves a target for a big acquihire.

Cynical take: Try to get acquired by Intel for Arc.

Intel is in a vastly better shape than AMD, they have the software pretty much nailed down.

Re: Making AMD GPUs competitive for LLM inference (2023)

#83
post #16
post #9

I have come across quite few startups who are trying a similar idea: break the nvidia monopoly by utilizing AMD GPUs (for inference at least): Felafax, Lamini, tensorwave (partially), SlashML. Even saw optimistic claims like CUDA moat is only 18 months deep from some of them [1]. Let's see. [1] https://www.linkedin.com/feed/update/urn:li:activity:7275885...

Tinygrad was another one, but they ended up getting frustrated with AMD and semi-pivoted to Nvidia.

> Tinygrad was another one, but they ended up getting frustrated with AMD and semi-pivoted to Nvidia.

From their announcement on 20241219[^0]:

"We are the only company to get AMD on MLPerf, and we have a completely custom driver that's 50x simpler than the stock one. A bit shocked by how little AMD cared, but we'll take the trillions instead of them."

From 20241211[^1]:

"We gave up and soon tinygrad will depend on 0 AMD code except what's required by code signing.

We did this for the 7900XTX (tinybox red). If AMD was thinking strategically, they'd be begging us to take some free MI300s to add support for it."

---

[^0]: https://x.com/__tinygrad__/status/1869620002015572023

[^1]: https://x.com/__tinygrad__/status/1866889544299319606

Re: Making AMD GPUs competitive for LLM inference (2023)

#84
post #23

Earlier quoted context omitted.

Is there no hope for AMD anymore? After George Hotz/Tinygrad gave up on AMD I feel there’s no realistic chance of using their chips to break the CUDA dominance.

Not really. AMD is constitutionally incapable of shipping anything but mid range hardware that requires no innovation. The only reason why they are doing so well in CPUs right now is that Intel has basically destroyed itself without any outside help.

It had to destroy itself. These companies do not act on their own...

Re: Making AMD GPUs competitive for LLM inference (2023)

#85
post #56

Earlier quoted context omitted.

And I'm supposed to believe that HN is this amazing platform for technology and science discussions, totally unlike its peers...

I don't really care what you believe. Everyone whose dug deep into what AMD is doing has left in disgust if they are lucky and bankruptcy if they are not. If I can save someone else from wasting $100,000 on hardware and six months of their life then my post has done more good than the AMD marketing department ever will.

> If I can save someone else from wasting $100,000 on hardware and six months of their life then my post has done more good than the AMD marketing department ever will.

This seems like unuseful advice if you've already given up on them.

You tried it and at some point in the past it wasn't ready. But by not being ready they're losing money, so they have a direct incentive to fix it. Which would take a certain amount of time, but once you've given up you no longer know if they've done it yet or not, at which point your advice would be stale.

Meanwhile the people who attempt it apparently seem to get acquired by Nvidia, for some strange reason. Which implies it should be a worthwhile thing to do. If they've fixed it by now which you wouldn't know if you've stopped looking, or they fix it in the near future, you have a competitive advantage because you have access to lower cost GPUs than your rivals. If not, but you've demonstrated a serious attempt to fix it for everyone yourself, Nvidia comes to you with a sack full of money to make sure you don't finish, and then you get a sack full of money. That's win/win, so rather than nobody doing it, it seems like everybody should be doing it.

Re: Making AMD GPUs competitive for LLM inference (2023)

#86

Earlier quoted context omitted.

They've actually been making positive moves with GPUs lately along with a success story for the B580.

Yeah but MLID says they are losing money on every one and have been winding down the internal development resources. That doesn't bode well for the future. I want to believe he's wrong, but on the parts of his show where I am in a position to verify, he generally checks out. Whatever the opposite of Gell-Mann Amnesia is, he's got it going for him.

The die size of the B580 is 272 mm2, which is a lot of silicon for $249. The performance of the GPU is good for its price but bad for its die size. Manufacturing cost is closely tied to die size.

272 mm2 puts the B580 in the same league as the Radeon 7700XT, a $449 card, and the GeForce 4070 Super, which is $599. The idea that Intel is selling these cards at a loss sounds reasonable to me.

Re: Making AMD GPUs competitive for LLM inference (2023)

#87

Great, I have yet to understand why does not the ML community really push or move away from CUDA? To me, it feel like a dinosaur move to build on top of CUDA which is screaming proprietary nothing about it is open source or cross platform. The reason why I say its dinosaur is, imagine, we as a dev community continued to build on top of Flash or Microsoft Silverlight... LLM and ML has been out for quiet a while, with…

Except I never hear complaints about CUDA from a quality perspective. The complaints are always about lock in to the best GPUs on the market. The desire to shift away is to make cheaper hardware with inferior software quality more usable. Flash was an abomination, CUDA is not.

Flash was popular because it was an attractive platform for the developer. Back then there was no HTML5 and browsers didn't otherwise support a lot of the things Flash did. Flash Player was an abomination, it was crashy and full of security vulnerabilities, but that was a problem for the user rather than the developer and it was the developer choosing what to use to make the site.

This is pretty much exactly what happens with CUDA. Developers like it but then the users have to use expensive hardware with proprietary drivers/firmware, which is the relevant abomination. But users have some ability to influence developers, so as soon as we get the GPU equivalent of HTML5, what happens?

Re: Making AMD GPUs competitive for LLM inference (2023)

#88
post #17

Earlier quoted context omitted.

Well sure, but in other GPU tasks, like Raytracing, the difference between these GPUs is far more pronounced. And AMD has passable Raytracing units (NVidias are better but the difference is bigger than these LLM results). If RAM is the main bottleneck then CPUs should be on the table.

> If RAM is the main bottleneck then CPUs should be on the table That's certainly not the case. The graphics memory model is very different from the CPU memory model. Graphics memory is explicitly designed for multiple simultaneous reads (spread across several different buses) at the cost of generality (only portions of memory may be available on each bus) and speed (the extra complexity means reads are slower). This…

> CPU memory only has one bus

If people are paying $15,000 or more per GPU, then I can choose $15,000 CPUs like EPYC that have 12-channels or dual-socket 24-channel RAM.

Even desktop CPUs are dual-channel at a minimum, and arguably DDR5 is closer to 2 or 4 buses per channel.

Now yes, GPU RAM can be faster, but guess what?

https://www.tomshardware.com/pc-components/cpus/amd-crafts-c...

GPUs are about extremely parallel performance, above and beyond what traditional single-threaded (or limited-SIMD) CPUs can do.

But if you're waiting on RAM anyway?? Then the compute-method doesn't matter. Its all about RAM.

Re: Making AMD GPUs competitive for LLM inference (2023)

#89
post #40

Earlier quoted context omitted.

Maybe from Modular (the company Chris Lattner is working for). In this recent announcement they said they had achieved competitive ML performance… on NVIDIA GPUs, but with their own custom stack completely replacing CUDA. And they’re targeting AMD next. https://www.modular.com/blog/introducing-max-24-6-a-gpu-nati...

Ah yes, the programming language (Mojo) that requires an account before I can use it...

Mojo no longer requires an account to install.

But that is irrelevant to the conversation because this is not about Mojo but something they call MAX. [1]

1. https://www.modular.com/max

Re: Making AMD GPUs competitive for LLM inference (2023)

#90
I got a "gaming" PC for LLM inference with an RTX 3060. I could have gotten more VRAM for my buck with AMD, but didn't because at the time alot of inference needed CUDA.

As soon AMD is as good as Nvidia for inference, I'll switch over.

But I've read on here that their hardware engineers aren't even given enough hardware to test with...

Post reply on HN