Live data from Hacker News

AMD's MI300X Outperforms Nvidia's H100 for LLM Inference

blog.tensorwave.com

141–150 of 273 posts

Re: AMD's MI300X Outperforms Nvidia's H100 for LLM Inference

#141
post #110
post #4

"TensorWave is a cloud provider specializing in AI workloads. Their platform leverages AMD’s Instinct™ MI300X accelerators, designed to deliver high performance for generative AI workloads and HPC applications." I suggest taking the report with a grain of salt.

The salt is in the plain sight. The do the standard AMD comparison: 8x AMD MI300X (192GB, 750W) GPU 8x H100 SXM5 (80GB, 700W) GPU The fair comparison would be against 8x H100 NVL (188GB, Price tells a story. If AMD performance would be in par with Nvidia they would not sell their cards for 1/4 price.

Isn't SXM5 higher bandwidth? It's 900 GB/s of bidirectional bandwidth per GPU across 18 NVLink 4 channels. The NVL's are on PCIe 5, and even w/ NVLink only get to 600 GB/s of bandwidth across 3 NVLink bridges (across only pairs of cards)?

I haven't done a head to head and I suppose it depends on whether tensor parallelism actually scales linearly or not, but my understanding is since the NVL's are just PCIe/NVLink paired H100s, you're not really getting much if any benefit on something like vLLM.

I think the more interesting thing critique might be the slightly odd choice of Mixtral 8x7B vs say a more standard Llama2/3 70B (or just test multiple models including some big ones like 8x22B or DBRX.

Also, while I don't have a problem w/ vLLM, as TensorRT gets easier to set up, it might become a factor in comparisons (since they punted on FP8/AMP in this tests). Inferless published a shootoff a couple months ago comparing a few different inference engines: https://www.inferless.com/learn/exploring-llms-speed-benchma...

Price/perf does tell a story, but I think it's one that's mostly about Nvidia's platform dominance and profit margins more than intrinsic hardware advantages. On the spec sheet MI300X has a memory bandwidth and even raw FLOPS advantage but so far it has lacked proper software optimization/support and wide availability (has anyone besides hyperscalers and select partners been able to get them?)

Re: AMD's MI300X Outperforms Nvidia's H100 for LLM Inference

#142
post #59

Earlier quoted context omitted.

Because tech innovation requires tons of R&D and you can't afford to do that otherwise. Europeans use American laptops running an American operating system to watch American movies in an American browser. European economic production is nowhere near high enough and now Europe is struggling to provide for its aging population and doesn't have enough good jobs for younger people. I support redistribution generally, but…

And yet standards of living in Europe are comparable to those in the US, and preferable at the median. Our attention is captured by speculative valuations of unicorns, and yet people actually need real stuff made, drugs developed and made etc. Europe does perfectly well in many non winner takes all sectors where English language and network effects are less relevant. The political instability created by the US neolib…

Europe is much much poorer than the US, actually.

https://pbs.twimg.com/media/F3PGpsrWEAEiplB?format=jpg&name=...

Re: AMD's MI300X Outperforms Nvidia's H100 for LLM Inference

#143

Earlier quoted context omitted.

It doesn't matter. AMD has offered better compute per dollar for a while now, but noone switched because CUDA is the real reason why all serious ML people use Nvidia. Until AMD picks up the slack on their software side, Nvidia will continue to dominate.

And this shouldn't be to hard if you know the ins and outs of the hardware and have a reasonable dev team. So why aren't they doing it?

> and have a reasonable dev team

probably this.

Re: AMD's MI300X Outperforms Nvidia's H100 for LLM Inference

#144
post #72

Earlier quoted context omitted.

The DAX is made of the 40 most valuable German companies. That’s how it is defined. So the companies not in it, again by definition, matter less.

> The DAX is made of the 40 most valuable German companies. Not to be too nitpicky here but these are only the publicly traded companies. You have a number of pretty large German companies that are still entirely private such as Aldi, Schwarz Group, Boehringer or Bosch.

Also many medium sized companies which are productive and competitive but not public, so never grow to huge sizes. I view this as a feature not a bug… smaller companies have a more direct connection with their workforce and tend to behave better with them.

Re: AMD's MI300X Outperforms Nvidia's H100 for LLM Inference

#145
post #104

Earlier quoted context omitted.

You can rent them online for ~ 4-5 $ per hour per GPU. Not cheap, but definitely feasible as a weekend project.

where can I rent a H100 for 4-5 dollars an hour? AWS doesn't let you use p5 instances (not getting a quota as a private person), lambda cloud is sold out.

It looks like Runpod currently (checked right now) has "Low" availability of 8x MI300 SXM (8x$4.89/h), H100 NVL (8x$4.39/h), and H100 (8x$4.69/h) nodes for anyone w/ some time to kill that wants to give the shootout a try.

Re: AMD's MI300X Outperforms Nvidia's H100 for LLM Inference

#146
post #110
post #4

"TensorWave is a cloud provider specializing in AI workloads. Their platform leverages AMD’s Instinct™ MI300X accelerators, designed to deliver high performance for generative AI workloads and HPC applications." I suggest taking the report with a grain of salt.

The salt is in the plain sight. The do the standard AMD comparison: 8x AMD MI300X (192GB, 750W) GPU 8x H100 SXM5 (80GB, 700W) GPU The fair comparison would be against 8x H100 NVL (188GB, Price tells a story. If AMD performance would be in par with Nvidia they would not sell their cards for 1/4 price.

Thx! Anyone who says Nivida isnt king, needs a reality check.

Re: AMD's MI300X Outperforms Nvidia's H100 for LLM Inference

#147
post #59

Earlier quoted context omitted.

Because tech innovation requires tons of R&D and you can't afford to do that otherwise. Europeans use American laptops running an American operating system to watch American movies in an American browser. European economic production is nowhere near high enough and now Europe is struggling to provide for its aging population and doesn't have enough good jobs for younger people. I support redistribution generally, but…

To be fair more than half of that laptop hardware is made in China+Taiwan, including a lot of the IP that goes into it. If you look at phones and tablets, there is a bunch of components/IP from European companies also, such as ARM, Bosch, STmicroelectronics, Infineon, NXP etc. Intel has famously struggled and failed multiple times to get into that market. European semiconductor companies are also strong in automotive…

That's all true. However, the US is capable of producing almost everything domestically, albeit in lower volume. The US military doesn't like to depend on Chinese chips for obvious reasons.

Re: AMD's MI300X Outperforms Nvidia's H100 for LLM Inference

#148

Earlier quoted context omitted.

> Making real physical things just doesn't scale, Nvidia sells chips ...

They can't make them fast enough to satisfy hype-driven demand; they are not scaling like a tech company, but their market cap is being inflated like one.

What is the meaning of "tech" in this context? "Software"? I mean, if Nvidia isn't a technology company...

Re: AMD's MI300X Outperforms Nvidia's H100 for LLM Inference

#149
post #141
post #110

Earlier quoted context omitted.

The salt is in the plain sight. The do the standard AMD comparison: 8x AMD MI300X (192GB, 750W) GPU 8x H100 SXM5 (80GB, 700W) GPU The fair comparison would be against 8x H100 NVL (188GB, Price tells a story. If AMD performance would be in par with Nvidia they would not sell their cards for 1/4 price.

Isn't SXM5 higher bandwidth? It's 900 GB/s of bidirectional bandwidth per GPU across 18 NVLink 4 channels. The NVL's are on PCIe 5, and even w/ NVLink only get to 600 GB/s of bandwidth across 3 NVLink bridges (across only pairs of cards)? I haven't done a head to head and I suppose it depends on whether tensor parallelism actually scales linearly or not, but my understanding is since the NVL's are just PCIe/NVLink pa…

> but I think it's one that's mostly about Nvidia's platform dominance and profit margins more

Profit margins and dominance are result from performance, not the other way around.

It does not matter if Nvidia tools are better when you deploy large number of chips for inference and it does more flops per watt or second. It's seller market and if AMD can't ask high price, their chip do not perform.

----

Question:

People here seem to think that Nvidia has absolutely no advantage in their microarchitecture design skills. It's all in software or monopoly.

Is this right?

Re: AMD's MI300X Outperforms Nvidia's H100 for LLM Inference

#150
post #13

I try to be optimistic about this. Competition is absolutely needed in this space - $NVDA market cap is insane right now, about $0.6 trillion more than the entire Frankfurt Stock Exchange.

We are in the middle of an LLM bubble. Nvidia problem will sort itself out naturally in the coming months/years.

As someone put it in: we are in the 3D glasses phase of AI. Remember when all TVs came with one?
Post reply on HN