Live data from Hacker News

AMD's MI300X Outperforms Nvidia's H100 for LLM Inference

blog.tensorwave.com

91–100 of 273 posts

Re: AMD's MI300X Outperforms Nvidia's H100 for LLM Inference

#92

Why the hell are we doing 128 input token benchmarks in 2024. This is not representative of most workloads, and prefill perf is incredibly important.

For understanding: What would be a suitable input length in your oppinion? And why isnt this a good one: Are real-life queries shorter? Or longer? If i count one word as a token, then in my case most of the queries are less than 128 words.

I think today 512 tokens is a minimum.

It's not just the query (if you're running a chatbot, which many of us are not). It's the entire context window. It's not uncommon to have a system prompt that is > 512 tokens alone.

I would like to see benchmarks for 512, 1024, 4096 and 8192 token inputs.

Re: AMD's MI300X Outperforms Nvidia's H100 for LLM Inference

#94
post #40

Earlier quoted context omitted.

Okay, so you are saying I should move to america, where apparently a lot of people struggle hard to even get a job? Nah, then ill get my very good wagie pennies here and have plenty jobs available, plus good health insurrance and whatnot.

German unemployment is 3.2%. US unemployment is 4.0%. Neither of these are at all high by historical standards. https://www.bls.gov/news.release/empsit.nr0.htm https://www.destatis.de/EN/Press/2024/06/PE24_217_132.html

Don't forget to take both of those stats with a grain of salt though. The US has a lot of gig workers which are not always counted correctly or the same and Germany has a large low-wage sector, where people are employed but earn less per month that they would get in unemployment benefits, and so the state pays the difference.

Re: AMD's MI300X Outperforms Nvidia's H100 for LLM Inference

#95

Earlier quoted context omitted.

It doesn't matter. AMD has offered better compute per dollar for a while now, but noone switched because CUDA is the real reason why all serious ML people use Nvidia. Until AMD picks up the slack on their software side, Nvidia will continue to dominate.

Large corporate customers like Microsoft and Meta do not use CUDA. They all use custom software. AMD doesn’t have enough GPUs to sell them yet, that’s the real bottleneck.

That's a pretty big claim, that Microsoft and Meta have their own proprietary cuda-replacement stack. Do you have any evidence for that claim?

Re: AMD's MI300X Outperforms Nvidia's H100 for LLM Inference

#96
post #4

"TensorWave is a cloud provider specializing in AI workloads. Their platform leverages AMD’s Instinct™ MI300X accelerators, designed to deliver high performance for generative AI workloads and HPC applications." I suggest taking the report with a grain of salt.

Well, there's the beauty of specifying exactly how you ran your benchmark, it is easy to reproduce and disprove or confirm (assuming you got the hardware).

As easy as getting yourself 8 H100 and 8 MI300X.

Fun weekend project for anybody.

Re: AMD's MI300X Outperforms Nvidia's H100 for LLM Inference

#97
post #39

Earlier quoted context omitted.

> The DAX is only 40 companies, most of which make real products rather than advertising mechanisms This, as the kids say, is just cope. American big tech makes real products. Google is not just ads. Apple is not. Amazon is not. Tesla is not. NVidia is not. Netflix is not. NVidia might be overvalued because of the current AI hype but that does not diminish their real accomplishments! Europe has almost no real tech co…

> There is one exception, founded in 1984 Europe clearly has many problems stimulating investments and creating a competitive environment for startups and tech companies. But you can't say ASML is the only one "real" tech company. What about Adyen, Spotify, Klarna, N26, Revolut, etc?

Most of those companies you mentioned aren't anywhere near as wealthy or as high market caps as US big-tech.

Most of them are just payment middlemen not some innovative product nobody else can do, and Spotify survives on monopolizing and squeezing artists, not some innovative product. Kind of like Netflix except Netflix has some cutting edge streaming tech as a product not just IP licenses.

ASML is the only product innovator there except their innovative EUV lightsources are licensed from Sandia labs in the US and made by Cymer in the US which ASML bought and licensed to not seel to China. So an US invention at the end of the day.

Re: AMD's MI300X Outperforms Nvidia's H100 for LLM Inference

#98

Earlier quoted context omitted.

If they used Nvidia's chip would this somehow make the blog post better?

For one, they didn't use TensorRT in the test. Also, stuff like this is hard to take the results seriously: * To make an accurate comparison between the systems with different settings of tensor parallelism, we extrapolate throughput for the MI300X by 2. * All inference frameworks are configured to use FP16 compute paths. Enabling FP8 compute is left for future work. They did everything they can to make sure AMD is f…

So they just multipiled their results per 2 ^^ ?

Re: AMD's MI300X Outperforms Nvidia's H100 for LLM Inference

#99

AMD has better seemingly better hardware - but not the production capacity to compete with Nvidia yet. Will be interesting to see margins compress when real competition catches up. Everybody thinks it’s CUDA that makes Nvidia the dominant player. It’s not - almost 40% of their revenue this year comes from mega corporations that use their own custom stack to interact with GPUs. It’s only a matter of time before compet…

Can you explain the cuda-less stack a little more or provide a source?

Re: AMD's MI300X Outperforms Nvidia's H100 for LLM Inference

#100
post #39

Earlier quoted context omitted.

> The DAX is only 40 companies, most of which make real products rather than advertising mechanisms This, as the kids say, is just cope. American big tech makes real products. Google is not just ads. Apple is not. Amazon is not. Tesla is not. NVidia is not. Netflix is not. NVidia might be overvalued because of the current AI hype but that does not diminish their real accomplishments! Europe has almost no real tech co…

> There is one exception, founded in 1984 Europe clearly has many problems stimulating investments and creating a competitive environment for startups and tech companies. But you can't say ASML is the only one "real" tech company. What about Adyen, Spotify, Klarna, N26, Revolut, etc?

N26, Klarna and Revolut are not tech companies. Neobanks are still banks. And Klarna is just modern store credit cards. These have been around for decades.

Adyen is very underrated, and Spotify is definitely tech.

Stripe should be on the list. DeepMind at one point.

Post reply on HN