Live data from Hacker News

AMD's MI300X Outperforms Nvidia's H100 for LLM Inference

blog.tensorwave.com

211–220 of 273 posts

Re: AMD's MI300X Outperforms Nvidia's H100 for LLM Inference

#211

The market (and selling price) is reflecting the perceived value of nvidia's solution vs AMDs - comprehensively including tooling, software, TCO and managability. Also curious how many companies are dropping that much money on those kind of accelerators just to run 8x 7B param models in parallel... You're also talking about being able to train a 14B model on a single accelerator. I'd be curious to see how "full-accel…

mi300x production is ramping , in latest earning report lisa su said 1H2024 is production capped , 2h2024 have increased production ( and still have some to sell ), thanks probably to cowos and hbm3/(e?) supply improved

large orders for those accelerators are placed months ahead

meanwhile mi300x on microsoft are fully booked...

https://techcommunity.microsoft.com/t5/azure-high-performanc...

"Scalable AI infrastructure running the capable OpenAI models These VMs, and the software that powers them, were purpose-built for our own Azure AI services production workloads. We have already optimized the most capable natural language model in the world, GPT-4 Turbo, for these VMs. ND MI300X v5 VMs offer leading cost performance for popular OpenAI and open-source models."

Re: AMD's MI300X Outperforms Nvidia's H100 for LLM Inference

#212

Earlier quoted context omitted.

Maybe I'm a naive fanboy, but I would put my money on Apple catching Nvidia before AMD or Intel.

But Apple doesn't produce servers or server hardware.

Currently no, but the XServe was a product for a decade. And they have built an internal ML cloud, presumably with rack mountable hardware. The bigger issue for apple IMO is they ditched the server features of their OS and they're not going to sell a hypothetical M4ultra Xserve with linux.

Re: AMD's MI300X Outperforms Nvidia's H100 for LLM Inference

#213

Earlier quoted context omitted.

I love AMD, but my Nvidia stock position currently is much higher than AMD.

AMD is at a much higher PE ratio. Is the market expecting AMD to up its game in the GPU sector? Or is the market expecting a pullback in GPU demand due to possibility for non-GPU AI solutions becoming the frontier or for AI investment to slow down?

Why not both?

Re: AMD's MI300X Outperforms Nvidia's H100 for LLM Inference

#214
post #134

Earlier quoted context omitted.

It doesn't matter. AMD has offered better compute per dollar for a while now, but noone switched because CUDA is the real reason why all serious ML people use Nvidia. Until AMD picks up the slack on their software side, Nvidia will continue to dominate.

Microsoft recently announced that they run chatgpt 3.5 & 4 on mi300 on Azure and the price/performance is better. https://www.amd.com/en/newsroom/press-releases/2024-5-21-amd...

I've used ChatGPT on Azure. It sucks on so many levels, everything about it was clearly enforced by some bean counters who see X dollars for Y flops with zero regard for developers. So choosing AMD here would be about par for the course. There is a reason why everyone at the top is racing to buy Nvidia cards and pay the premium.

Re: AMD's MI300X Outperforms Nvidia's H100 for LLM Inference

#215
post #115

Earlier quoted context omitted.

I would argue that the EU is quite a successful organization given the task of setting up general market rules for 26 countries. Obviously there is a lot to criticize about the EU and I can offer you a gigantic list there too. However, I do not see any clear failure of the EU’s approach as a single market so far. Additionally part of the philosophy was establishing peace in a region that was torn up by wars for a lot…

It's definitely successful, and I was probably too harsh there. But I genuinely think the barriers that are left are damn near insurmountable. An awful lot has to change before a Greek tech workers can move to Sweden as easily as a Virginian can move to California.

I had two Greeks on my team at a large tech company in Sweden.

Re: AMD's MI300X Outperforms Nvidia's H100 for LLM Inference

#216
post #199

The market (and selling price) is reflecting the perceived value of nvidia's solution vs AMDs - comprehensively including tooling, software, TCO and managability. Also curious how many companies are dropping that much money on those kind of accelerators just to run 8x 7B param models in parallel... You're also talking about being able to train a 14B model on a single accelerator. I'd be curious to see how "full-accel…

the market and the selling price also includes sales strategies, penetrating a sector dominated by a strong player with somewhat "smart" sales strategies *1 and with a growing but certainly less mature product ( expecially software ), it requires suitable pricing and allocation strategies 1. https://www.techspot.com/news/102056-nvidia-allegedly-punish...

This stuff is the actual reason nvidia is under antitrust investigation.

boo boo, a GTX 670 that cost you $399 in 2012 now costs $599 - grow up, do the inflation calculation, and realize you’re being a child. gamers get the best deal on bulk silicon on the planet, R&D subsidized by enterprise, fantastic blue-sky research that takes years for competitors to (not even) match, and it’s still never enough. ”Gamers” have justified every single cliche and stereotype over the last 5 years, absolutely inveterate manbabies.

(Hardware Unboxed put out a video today with the headline+caption combo “are gamers entitled”/“are GeForce gpus gross”, and that’s what passes for reasoned discourse among the most popular channels. They’ve been trading segments back and forth with GN that are just absolute “how bad is nvidia” “real bad, but what do you guys think???” tier shit, lmao.

https://i.imgur.com/98x0F1H.png

this stuff is real shit, nvidia has been leaning on partners to maintain their segmentation, micromanaging shipment release to maintain price levels (cartel behavior), punishing customers and suppliers with “you know what will happen if you cross us”, literally putting it in writing with GPP (big mistake), playing fuck fuck games with not letting the drivers be run in a datacenter, etc. You see how that’s a little different than a gpu going from an inflation-adjusted $570 to $599 over 10 years?

(And what’s worse the competition can’t even keep that much, they’re falling off even harder now that Moores law has really kicked the bucket and they have to do architectural work every gen just to make progress, instead of getting free shrinks etc… let alone having to develop software! /gasp)

In entirely unrelated news… gigabyte suddenly has a 4070 ti super with a blower cooler. Oh, and it’s single-slot with end-fire power connector. All three forbidden features at once - very subtle, extremely law-abiding.

https://videocardz.com/newz/gigabyte-unveils-geforce-rtx-407...

and literally gamers can’t help but think this whole ftc case is all about themselves anyway…

Re: AMD's MI300X Outperforms Nvidia's H100 for LLM Inference

#217
post #160

Earlier quoted context omitted.

[flagged]

Even if a company doesn’t pay taxes, its masses of highly paid workers do. Each of those SWEs making $300k+ are paying more in taxes than the entire earnings of the average EU dev.

Aha, the cries of the once well-payed SWE laid off and replaced with some cheap hire oversea come up at least once a week for a year or two already. The funny thing about transnationals - they do not care about the society they operate in.

But what about the rest, not these lucky SWE who had a good run for the last 10-15 years? How is it going, education, medicine, crime, inequality? All is splendid, I assume?

Re: AMD's MI300X Outperforms Nvidia's H100 for LLM Inference

#218

Earlier quoted context omitted.

It doesn't matter. AMD has offered better compute per dollar for a while now, but noone switched because CUDA is the real reason why all serious ML people use Nvidia. Until AMD picks up the slack on their software side, Nvidia will continue to dominate.

Unless you develop in CUDA, you can easily train code (e.g. PyTorch) written for training on Nvidia hardware on AMD hardware. You can even keep the .cuda() calls.

In theory. But if you actually work with that in practice, you're already going to have a bad experience installing the drivers. And it's all downhill from there.

Re: AMD's MI300X Outperforms Nvidia's H100 for LLM Inference

#219
post #131

Earlier quoted context omitted.

MTr ------------------ H100 SXM5 80,000 MI300X 153,000 H100 NVL 160,000 H100 SXM4 has 52% of the transistors MI300X has, half of the RAM and MI300X achieves *ONLY* 33% higher throughput compared to the H100. MI300X was launched 6 months ago, H100 20 months ago. AMD has work to do.

Maybe I'm a naive fanboy, but I would put my money on Apple catching Nvidia before AMD or Intel.

Apple doesn't have any hardware SIMD technology that I'm aware of.

At best, Apple has Metal API which iOS video games use. I guess there's a level of SIMD-compute expertise here, but it'd take a lot of investment to turn that into a full scale GPU that tangos with Supercomputers. Software is a bit piece of the puzzle for sure, but Metal isn't ready for prime time.

I'd say Apple is ahead of Intel (Intel keeps wasting their time and collapsing their own progress from Xeon Phi / Battlemage / etc. etc. Intel cannot keep investing in its own stuff to reach critical mass). Intel does have OneAPI but given how many times Intel collapses everything and starts over again, I'm not sure how long OneAPI will last.

But Apple vs AMD? AMD 100% understands SIMD compute and has decades worth of investments in it. The only problem with AMD is that they don't have the raw cash to build out their expertise to cover software, so AMD has to rely upon Microsoft (DirectX), Vulkan, or whatever. ROCm may have its warts, but it does represent over a decade of software development too (especially when we consider that ROCm was "Boltzmann", which had several years of use before it came out as ROCm).

-------

AMD ain't perfect. They had a little diversion with C++Amp with Microsoft (and this served as the API for Boltzmann / early ROCm). But the overall path AMD is making at least makes sense, if a bit suboptimal compared to NVidia's huge efforts into CUDA.

I'd definitely rate AMD's efforts above Apple's Metal.

Re: AMD's MI300X Outperforms Nvidia's H100 for LLM Inference

#220

Earlier quoted context omitted.

LLVM IR to machine code is not the part that AMD has traditionally struggled with. What you call "trivial" is. If everyone started emitting IR and didn't rely on NVidia-owned libs then the space would become unrecognizable. The codegen is something AMD has always been decent at, hence them beating NVidia in compute benchmarks for most of the past 20 years.

> LLVM IR to machine code is not the part that AMD has traditionally struggled with. alright fine it's the codegen and the runtime and the driver and the library ecosystem... > If everyone started emitting IR and didn't rely on NVidia-owned libs then the space would become unrecognizable. I have no clue what this means - which libs are you talking about here? the libs that contain the implementations of their runtime…

AMD is struggling with unsafe C and C++ code breaking their drivers.
Post reply on HN