Live data from Hacker News

AMD's MI300X Outperforms Nvidia's H100 for LLM Inference

blog.tensorwave.com

161–170 of 273 posts

Re: AMD's MI300X Outperforms Nvidia's H100 for LLM Inference

#161
post #110

Earlier quoted context omitted.

The salt is in the plain sight. The do the standard AMD comparison: 8x AMD MI300X (192GB, 750W) GPU 8x H100 SXM5 (80GB, 700W) GPU The fair comparison would be against 8x H100 NVL (188GB, Price tells a story. If AMD performance would be in par with Nvidia they would not sell their cards for 1/4 price.

Thx! Anyone who says Nivida isnt king, needs a reality check.

I love AMD, but my Nvidia stock position currently is much higher than AMD.

Re: AMD's MI300X Outperforms Nvidia's H100 for LLM Inference

#162
post #142

Earlier quoted context omitted.

Europe is much much poorer than the US, actually. https://pbs.twimg.com/media/F3PGpsrWEAEiplB?format=jpg&name=...

EU GDP per capita 2022 is the same as US GDP per capita 2017. Unless you want to say that the US was much poorer in 2017 than it was in 2022 that's a fairly ridiculous statement. Also, the highest productivity places in the EU have much lower hours worked per capita than the US, with Germans on average working 25% less than Americans and the EU as a whole working 13% less than the US. https://data.oecd.org/emp/hours-…

Yes, if you had 2017 salary in the US today you are much poorer. Inflation was no joke

Re: AMD's MI300X Outperforms Nvidia's H100 for LLM Inference

#163
post #59

Earlier quoted context omitted.

Because tech innovation requires tons of R&D and you can't afford to do that otherwise. Europeans use American laptops running an American operating system to watch American movies in an American browser. European economic production is nowhere near high enough and now Europe is struggling to provide for its aging population and doesn't have enough good jobs for younger people. I support redistribution generally, but…

.... while all of the content is hosted on Linux servers, and you're listening to music through a Swedish app. That is running on silicone made in Taiwan. Using equipment that can currently only be manufactured in Belgium and Germany. On a Mac you are using a British instruction set. Also in terms of tech innovation: What part of the US-based tech innovation couldn't have been (and actually were) achieved with open-s…

>EU per Capita GDP in 2022 is the same as USA 2016. Was the USA in 2016 struggling but now isn't?

Absolutely the US would be struggling if the GDP was still 2016 values with today’s costs.

Re: AMD's MI300X Outperforms Nvidia's H100 for LLM Inference

#165
post #110

Earlier quoted context omitted.

The salt is in the plain sight. The do the standard AMD comparison: 8x AMD MI300X (192GB, 750W) GPU 8x H100 SXM5 (80GB, 700W) GPU The fair comparison would be against 8x H100 NVL (188GB, Price tells a story. If AMD performance would be in par with Nvidia they would not sell their cards for 1/4 price.

AMDs deep learning libraries are very bad the last time I checked, nobody uses amd in that space for that reason. Nvidia has a quazi monopoly, that's the main reason for the price difference IMHO.

this...

nearly 95% of deeplearning github repos are "tested using cuda gpu - others, not so sure"

the only way to out-run nvidia is to have 3~10x better bang-for-buck.

Or AMD can just provide a "DIY unlimited gpu RAM upgrade" kit -- a lot of people are buying macstudio 128gb ram because of its "bigger ram-for-buck" than nvidia gpus

Re: AMD's MI300X Outperforms Nvidia's H100 for LLM Inference

#166
post #34

Earlier quoted context omitted.

I see it as they did everything they can to compare the specific code path. If your workload scales with FP16 but not with tensor cores, then this is the correct way to test. What do you need for LLM inference?

Couldn't they find a real workload that does this?

vLLM inference of Mixtral in fp16 is a real workload. I guess the details are there because of the different inference engine used. You need the most similar compute tasks to be ran but the compute kernels can't be the same as in the end they need to be ran by a different hardware.

Re: AMD's MI300X Outperforms Nvidia's H100 for LLM Inference

#167

Earlier quoted context omitted.

Same thing was said about Nvidia's crypto bubbles, and then look what happened. Jensen isn't stupid. He's making accelerators for anything so that they'll be ready to catch the next bubble that depends on crazy compute power that can't be done efficiently on CPUs. They're so far the only semi company beating Moore's law by a large margin due to their clever scaling tech while everyone else is like "hey look our new p…

They got extremely lucky with AI following crypto. The timing was close to perfect. I'm not sure there will be another wave like that at all for a long while.

It wasn't a coincidence either though: the amount of compute available is probably the main driver of this wave.

Re: AMD's MI300X Outperforms Nvidia's H100 for LLM Inference

#168
post #59

Earlier quoted context omitted.

Because tech innovation requires tons of R&D and you can't afford to do that otherwise. Europeans use American laptops running an American operating system to watch American movies in an American browser. European economic production is nowhere near high enough and now Europe is struggling to provide for its aging population and doesn't have enough good jobs for younger people. I support redistribution generally, but…

To be fair more than half of that laptop hardware is made in China+Taiwan, including a lot of the IP that goes into it. If you look at phones and tablets, there is a bunch of components/IP from European companies also, such as ARM, Bosch, STmicroelectronics, Infineon, NXP etc. Intel has famously struggled and failed multiple times to get into that market. European semiconductor companies are also strong in automotive…

>If you look at phones and tablets, there is a bunch of components/IP from European companies also, such as ARM, Bosch, STmicroelectronics, Infineon, NXP etc.

But those are all low-marin chips. Qualcomm, Nvidia, Intel, AMD and Apple have much higher margins on their chips. They don't bother competing with the EU chips companies.

Re: AMD's MI300X Outperforms Nvidia's H100 for LLM Inference

#169

Earlier quoted context omitted.

If they used Nvidia's chip would this somehow make the blog post better?

For one, they didn't use TensorRT in the test. Also, stuff like this is hard to take the results seriously: * To make an accurate comparison between the systems with different settings of tensor parallelism, we extrapolate throughput for the MI300X by 2. * All inference frameworks are configured to use FP16 compute paths. Enabling FP8 compute is left for future work. They did everything they can to make sure AMD is f…

You need 2 H100 to have enough VRAM for the model whereas you need only 1 MI300X. Doubling the total throughput (for all completions) of 1 MI300X to simulate the numbers for a duplicated system is reasonable.

They should probably show separately the throughput per completion as the tensor parallelism is often used for that purpose in addition to the doubling the VRAM.

Re: AMD's MI300X Outperforms Nvidia's H100 for LLM Inference

#170
post #156

Earlier quoted context omitted.

EU GDP per capita 2022 is the same as US GDP per capita 2017. Unless you want to say that the US was much poorer in 2017 than it was in 2022 that's a fairly ridiculous statement. Also, the highest productivity places in the EU have much lower hours worked per capita than the US, with Germans on average working 25% less than Americans and the EU as a whole working 13% less than the US. https://data.oecd.org/emp/hours-…

World Bank: European Union gdp per capita for 2022 was $37,433, a 3.33% decline from 2021. U.S. gdp per capita for 2022 was $76,330, a 8.7% increase from 2021. It's not even close?

GDP per capita as a stand alone metric doesn't mean much.
Post reply on HN