Earlier quoted context omitted.
That's a pretty big claim, that Microsoft and Meta have their own proprietary cuda-replacement stack. Do you have any evidence for that claim?
I'm guessing what they meant is that they use toolchains that are retargetable to other GPUs (and typically compile down to PTX (nVidia assembly language) on nVidia GPUs rather than go through CUDA source -- GCC and clang can both target PTX). For example XLA and most SYSCL toolchains support much more than nVidia.
AMD's MI300X Outperforms Nvidia's H100 for LLM Inference
251–260 of 273 posts
Re: AMD's MI300X Outperforms Nvidia's H100 for LLM Inference
#252Earlier quoted context omitted.
I love AMD, but my Nvidia stock position currently is much higher than AMD.
AMD is at a much higher PE ratio. Is the market expecting AMD to up its game in the GPU sector? Or is the market expecting a pullback in GPU demand due to possibility for non-GPU AI solutions becoming the frontier or for AI investment to slow down?
AMD's PE is ~55. Nvidia's PE is above 70.
Re: AMD's MI300X Outperforms Nvidia's H100 for LLM Inference
#253Earlier quoted context omitted.
If they used Nvidia's chip would this somehow make the blog post better?
For one, they didn't use TensorRT in the test. Also, stuff like this is hard to take the results seriously: * To make an accurate comparison between the systems with different settings of tensor parallelism, we extrapolate throughput for the MI300X by 2. * All inference frameworks are configured to use FP16 compute paths. Enabling FP8 compute is left for future work. They did everything they can to make sure AMD is f…
Re: AMD's MI300X Outperforms Nvidia's H100 for LLM Inference
#254Earlier quoted context omitted.
These kinds of comments make me think few people have actually tried. My experience has been 1 work day of getting things set up to work the same as before for training and testing (PyTorch).
You have to consider that the average person who tried to do machine learning on AMD GPUs got burned in the past decade and has no reason to change their opinion. Also in the past it was much harder to get access to cutting edge GPUs from AMD. The fact that AMD drops GPU support for ROCm quickly also earns them scorn. I don't think it is an unfair assessment. They earned their reputation.
Re: AMD's MI300X Outperforms Nvidia's H100 for LLM Inference
#255Earlier quoted context omitted.
So? If you get twice the RAM at a comparable price and that leads to twice the performance, what's wrong with comparing that?
Nothing wrong - just for transparency. Also, the price difference is not quantified. Additionally, CUDA is a known and tangible software stack - can I try out this "MK1 FLywheel" on my local (AMD) hardware?
Re: AMD's MI300X Outperforms Nvidia's H100 for LLM Inference
#256Earlier quoted context omitted.
Blackwell won't be here till next year.
GB200 based on blackwell launched in March of this year. https://www.theregister.com/2024/03/21/nvidia_dgx_gb200_nvk7... MI300X launched 3 months earlier at the end of December. H100 launched March 2023,
Also Blackwell's lead will be short lived, because mi350x is coming out next year, and it will have a node and architecture advantage. So AMD will be ahead again.
Re: AMD's MI300X Outperforms Nvidia's H100 for LLM Inference
#257https://www.reddit.com/r/AMD_MI300/comments/1dgimxt/benchmar...
Re: AMD's MI300X Outperforms Nvidia's H100 for LLM Inference
#258Earlier quoted context omitted.
Microsoft recently announced that they run chatgpt 3.5 & 4 on mi300 on Azure and the price/performance is better. https://www.amd.com/en/newsroom/press-releases/2024-5-21-amd...
I've used ChatGPT on Azure. It sucks on so many levels, everything about it was clearly enforced by some bean counters who see X dollars for Y flops with zero regard for developers. So choosing AMD here would be about par for the course. There is a reason why everyone at the top is racing to buy Nvidia cards and pay the premium.
It looks like the price to performance of inference tasks gives providers a big incentive to move away from Nvidia.
Re: AMD's MI300X Outperforms Nvidia's H100 for LLM Inference
#259> Hardware: TensorWave node equipped with 8 MI300X accelerators, 2 AMD EPYC CPU Processors (192 cores), and 2.3 TB of DDR5 RAM. > MI300X Accelerator: 192GB VRAM, 5.3 TB/s, ~1300 TFLOPS for FP16 > Hardware: Baremetal node with 8 H100 SXM5 accelerators with NVLink, 160 CPU cores, and 1.2 TB of DDR5 RAM. > H100 SXM5 Accelerator: 80GB VRAM, 3.35 TB/s, ~986 TFLOPS for FP16 I really wonder about the pricing. In theory the…
RunPod [0] is pricing MI300X at $4.89/hr vs $3.89-4.69/hr for H100s. So, probably around the same price? The tests look promising, though! [0] https://runpod.io/
The weird thing on Runpod is the virtual CPUs, you can't run MI300x in virtual machines yet. It is a missing feature that AMD is working on.
Re: AMD's MI300X Outperforms Nvidia's H100 for LLM Inference
#260Is this an ad for a new, closed-source, GPGPU backend?
https://www.reddit.com/r/AMD_MI300/comments/1dgimxt/benchmar...