Live data from Hacker News

AMD's MI300X Outperforms Nvidia's H100 for LLM Inference

blog.tensorwave.com

241–250 of 273 posts

Re: AMD's MI300X Outperforms Nvidia's H100 for LLM Inference

#241

Earlier quoted context omitted.

AWS is a (collection of) service(s) which can scale and have low marginal cost of reproduction. Fire Tablets are a real thing you can buy, but they are shit so nobody does.

Sorry, are you suggesting AWS Services "aren't real things"? If I pay for a database server in Virginia, how is that not real ?

I get that you meant “physical” but this blur between “advertising isn’t a useful thing for an economy to focus on” and “services are not real” is a bit of a jump!

Cloud services don’t just exist on their own accord. Datacenters are physical and real!

Re: AMD's MI300X Outperforms Nvidia's H100 for LLM Inference

#242
post #110
post #4

"TensorWave is a cloud provider specializing in AI workloads. Their platform leverages AMD’s Instinct™ MI300X accelerators, designed to deliver high performance for generative AI workloads and HPC applications." I suggest taking the report with a grain of salt.

The salt is in the plain sight. The do the standard AMD comparison: 8x AMD MI300X (192GB, 750W) GPU 8x H100 SXM5 (80GB, 700W) GPU The fair comparison would be against 8x H100 NVL (188GB, Price tells a story. If AMD performance would be in par with Nvidia they would not sell their cards for 1/4 price.

fair? h100 NVL are two h100 in a single package.. which probably costs 2xh100 or more,

if so ok it's fair to compare 1 mi300x with 1 h100 NVL but then price ( and tco ) should be added to the some metrics conclusion , also the NVL is a 2xpci5.0 quad slot , so not the same thing..

I am not sure about system compatibility and if and how you can stack 8 of those in one system ( like you can do with non NVL and mi300x.. ) so it's a bit a diffent ( and more niche ) beast

Re: AMD's MI300X Outperforms Nvidia's H100 for LLM Inference

#243

Earlier quoted context omitted.

Apple doesn't have any hardware SIMD technology that I'm aware of. At best, Apple has Metal API which iOS video games use. I guess there's a level of SIMD-compute expertise here, but it'd take a lot of investment to turn that into a full scale GPU that tangos with Supercomputers. Software is a bit piece of the puzzle for sure, but Metal isn't ready for prime time. I'd say Apple is ahead of Intel (Intel keeps wasting…

Apple makes a better consumer GPU than AMD does. M3 Max's GPU is significantly more efficient in perf/watt than RDNA3, already has better ray tracing performance, and is even faster than a 7900XT desktop GPU in Blender.[0] [0] https://opendata.blender.org/benchmarks/query/?compute_type=...

Couple of things: Blender uses HIP for AMD which is nerfed in RDNA3 because of product segmentation, so really this is comparing against something which is deliberately mediocre in the 7900 XT.

The M3 Max is also in a sense a generation ahead in terms of perf/watt of the 7900 XT as it uses a newer manufacturing node.

I suppose it's also worth highlighting that if you enable Optix in the comparison above, you can see Nvidia parts stomping all over both AMD and Apple parts alike.

Re: AMD's MI300X Outperforms Nvidia's H100 for LLM Inference

#244

Earlier quoted context omitted.

Apple makes a better consumer GPU than AMD does. M3 Max's GPU is significantly more efficient in perf/watt than RDNA3, already has better ray tracing performance, and is even faster than a 7900XT desktop GPU in Blender.[0] [0] https://opendata.blender.org/benchmarks/query/?compute_type=...

Couple of things: Blender uses HIP for AMD which is nerfed in RDNA3 because of product segmentation, so really this is comparing against something which is deliberately mediocre in the 7900 XT. The M3 Max is also in a sense a generation ahead in terms of perf/watt of the 7900 XT as it uses a newer manufacturing node. I suppose it's also worth highlighting that if you enable Optix in the comparison above, you can see…

Why does AMD nerf RDNA3 when they're so far behind Nvidia and Apple in Blender performance? Do you have benchmarks for when AMD doesn't nerf Blender performance?

M3 Max GPU uses at most 60-70w. Meanwhile, the 7900XT uses up to 412w in burst mode.[0] TSMC N3 (M3 Max) uses 25-30% less power than TSMC N5 (7900XT). [1] In other words, if 7900XT used N3 and optimizes for the same performance, it would burst to 300w instead which is still 5-6x more than M3 Max. In other words, the perf/watt advantage of the M3 Max is mostly not related to the node used. It's the design.

[0]https://www.techpowerup.com/review/amd-radeon-rx-7900-xt/37....

[1]https://www.anandtech.com/show/18833/tsmc-details-3nm-evolut...

Re: AMD's MI300X Outperforms Nvidia's H100 for LLM Inference

#245
post #110
post #4

"TensorWave is a cloud provider specializing in AI workloads. Their platform leverages AMD’s Instinct™ MI300X accelerators, designed to deliver high performance for generative AI workloads and HPC applications." I suggest taking the report with a grain of salt.

The salt is in the plain sight. The do the standard AMD comparison: 8x AMD MI300X (192GB, 750W) GPU 8x H100 SXM5 (80GB, 700W) GPU The fair comparison would be against 8x H100 NVL (188GB, Price tells a story. If AMD performance would be in par with Nvidia they would not sell their cards for 1/4 price.

Price tells the story. Yes but for electric prices not card prize and here their much more close to each other!

Re: AMD's MI300X Outperforms Nvidia's H100 for LLM Inference

#246

Earlier quoted context omitted.

Couple of things: Blender uses HIP for AMD which is nerfed in RDNA3 because of product segmentation, so really this is comparing against something which is deliberately mediocre in the 7900 XT. The M3 Max is also in a sense a generation ahead in terms of perf/watt of the 7900 XT as it uses a newer manufacturing node. I suppose it's also worth highlighting that if you enable Optix in the comparison above, you can see…

Why does AMD nerf RDNA3 when they're so far behind Nvidia and Apple in Blender performance? Do you have benchmarks for when AMD doesn't nerf Blender performance? M3 Max GPU uses at most 60-70w. Meanwhile, the 7900XT uses up to 412w in burst mode.[0] TSMC N3 (M3 Max) uses 25-30% less power than TSMC N5 (7900XT). [1] In other words, if 7900XT used N3 and optimizes for the same performance, it would burst to 300w instea…

Its weird that you're choosing a nerf'd part and sticking with it as a comparison point.

The article is MI300X, which is beating NVidia's H100.

> Do you have benchmarks for when AMD doesn't nerf Blender performance?

Go read the article above.

> Notably, our results show that MI300X running MK1 Flywheel outperforms H100 running vLLM for every batch size, with an increase in performance ranging from 1.22x to 2.94x.

-------

> Why does AMD nerf RDNA3 when they're so far behind Nvidia and Apple in Blender performance?

Nerf is a weird word.

AMD has focused on 32-bit FLOPs and 64-bit FLOPs until now. AMD never put much effort into raytracing. They reach acceptable levels on XBox / PS5 but NVidia always was pushing Raytracing (not AMD).

Similarly: Blender is a raytracer that uses those Raytracing cores. So any chip with substantial on-chip ray-tracing / ray-matching / ray-intersection routines will perform faster.

Blender isn't what people do with GPUs. The #1 thing they do is video games like Baldur's gate 3.

-------

It'd be like me asking why Apple's M3 can't run Baldur's gate 3. Its not a "nerf", its a purposeful engineering decision.

Re: AMD's MI300X Outperforms Nvidia's H100 for LLM Inference

#247

Earlier quoted context omitted.

Maybe I'm a naive fanboy, but I would put my money on Apple catching Nvidia before AMD or Intel.

But Apple doesn't produce servers or server hardware.

Clearly they are building their own Apple Silicon powered servers for their Private Computing Cloud, even if it is not sold to outsiders like the XServe used to be.

Re: AMD's MI300X Outperforms Nvidia's H100 for LLM Inference

#248

Earlier quoted context omitted.

> but not the production capacity to compete with Nvidia yet. thats just a question of negotiating with tsmc or their few competitors (also didn't tsmc start production of some factories in the US and/or EU?) I mean, nvidia use tsmc, so does amd.

Yes it is - but Nvidia has larger contracts _right now_. Nvidia has been investing more money in producing more GPUs for longer, so it’s only natural that they have an advantage now. But now that there’s a larger incentive to produce GPUs, their moat will eventually fall. TSMC runs at 100% capacity for top tier processes - their bottleneck is more foundries. These take time to build. So the question becomes - how lon…

Isn't their moat primarily software (CUDA) rather than supply-chain strength?

Re: AMD's MI300X Outperforms Nvidia's H100 for LLM Inference

#249

Earlier quoted context omitted.

Why does AMD nerf RDNA3 when they're so far behind Nvidia and Apple in Blender performance? Do you have benchmarks for when AMD doesn't nerf Blender performance? M3 Max GPU uses at most 60-70w. Meanwhile, the 7900XT uses up to 412w in burst mode.[0] TSMC N3 (M3 Max) uses 25-30% less power than TSMC N5 (7900XT). [1] In other words, if 7900XT used N3 and optimizes for the same performance, it would burst to 300w instea…

Its weird that you're choosing a nerf'd part and sticking with it as a comparison point. The article is MI300X, which is beating NVidia's H100. > Do you have benchmarks for when AMD doesn't nerf Blender performance? Go read the article above. > Notably, our results show that MI300X running MK1 Flywheel outperforms H100 running vLLM for every batch size, with an increase in performance ranging from 1.22x to 2.94x. ---…

I was responding to the person above me, who used the word "nerf" to describe RDNA and Blender.

Re: AMD's MI300X Outperforms Nvidia's H100 for LLM Inference

#250
post #169

Earlier quoted context omitted.

For one, they didn't use TensorRT in the test. Also, stuff like this is hard to take the results seriously: * To make an accurate comparison between the systems with different settings of tensor parallelism, we extrapolate throughput for the MI300X by 2. * All inference frameworks are configured to use FP16 compute paths. Enabling FP8 compute is left for future work. They did everything they can to make sure AMD is f…

You need 2 H100 to have enough VRAM for the model whereas you need only 1 MI300X. Doubling the total throughput (for all completions) of 1 MI300X to simulate the numbers for a duplicated system is reasonable. They should probably show separately the throughput per completion as the tensor parallelism is often used for that purpose in addition to the doubling the VRAM.

What's the cost to run 2x H100 and 1x MI300X?

I think that'd give us a better idea of perf/cost and whether multiplying MI300X results by 2 is justified.

Post reply on HN