Live data from Hacker News

AMD's MI300X Outperforms Nvidia's H100 for LLM Inference

blog.tensorwave.com

261–270 of 273 posts

Re: AMD's MI300X Outperforms Nvidia's H100 for LLM Inference

#261

Earlier quoted context omitted.

I'm guessing what they meant is that they use toolchains that are retargetable to other GPUs (and typically compile down to PTX (nVidia assembly language) on nVidia GPUs rather than go through CUDA source -- GCC and clang can both target PTX). For example XLA and most SYSCL toolchains support much more than nVidia.

Even then it's an insanely bold assumption that a company other than Nvidia could build a better framework than CUDA for compiling PTX. Especially since CUDA is already so much C-like. I've never seen anyone go deeper than that outside of academia.

How many customers/consumers will care about services are be built with CUDA?

If they need a ChatBot that uses a model with same accuracy and performance as on non-CUDA hardware, would they still want CUDA based hardware?

Re: AMD's MI300X Outperforms Nvidia's H100 for LLM Inference

#262
post #197

Earlier quoted context omitted.

mi300x win in some inference workloads, h100 win in training and some others inference workloads ( fp8 inference with tensorRT-llm , rocm is young but is growing fast ) in a single system ( 8x accelerators ) LLMs, mi300x has very competitive inference TCO vs h100 . also : AMD Instinct MI300X Offers The Best Price To Performance on GPT-4 According To Microsoft, Red Team On-Track For 100x Perf/Watt By 2027 https://wccf…

wccftech is an untrustworthy source.

the article contain a quote from satya nadella : https://x.com/ryanshrout/status/1792953227841015897

Re: AMD's MI300X Outperforms Nvidia's H100 for LLM Inference

#263
post #262

Earlier quoted context omitted.

wccftech is an untrustworthy source.

the article contain a quote from satya nadella : https://x.com/ryanshrout/status/1792953227841015897

None of the mentioned claims from the article are confirmed there.

Re: AMD's MI300X Outperforms Nvidia's H100 for LLM Inference

#264
post #145

Earlier quoted context omitted.

It looks like Runpod currently (checked right now) has "Low" availability of 8x MI300 SXM (8x$4.89/h), H100 NVL (8x$4.39/h), and H100 (8x$4.69/h) nodes for anyone w/ some time to kill that wants to give the shootout a try.

We'd be happy to provide access to MI300X at TensorWave so you can validate our results! Just shoot us an email or fill out the form on our website

If you're able to advertise available GPU compute in some public forums then it's enough to tell us about the demand of MI300X in cloud ...

Re: AMD's MI300X Outperforms Nvidia's H100 for LLM Inference

#265

Earlier quoted context omitted.

Sorry, are you suggesting AWS Services "aren't real things"? If I pay for a database server in Virginia, how is that not real ?

I get that you meant “physical” but this blur between “advertising isn’t a useful thing for an economy to focus on” and “services are not real” is a bit of a jump! Cloud services don’t just exist on their own accord. Datacenters are physical and real!

fortunately, it's not for a one person to decide what an economy needs or not. The demand is there and that's why we have "ad companies".

Re: AMD's MI300X Outperforms Nvidia's H100 for LLM Inference

#266

Earlier quoted context omitted.

Same thing was said about Nvidia's crypto bubbles, and then look what happened. Jensen isn't stupid. He's making accelerators for anything so that they'll be ready to catch the next bubble that depends on crazy compute power that can't be done efficiently on CPUs. They're so far the only semi company beating Moore's law by a large margin due to their clever scaling tech while everyone else is like "hey look our new p…

I think there's also a very high prospect of virtual worlds with virtual people (SFW or otherwise) becoming popular, rendered with Apple/META goggles...that could require insane amounts of compute. And this is just one possibility. Relatively cheap multimodal smart glasses you wear when out and around that offload compute to the cloud are another. Nvidia could just as easily triple in short order as get cut in half f…

You mean like Omniverse from Nvidia where you can simulate entire factories or as Nvidia does data centers before they are built?

Or how about building the world virtually? https://www.nvidia.com/en-us/high-performance-computing/eart...

Re: AMD's MI300X Outperforms Nvidia's H100 for LLM Inference

#267

Earlier quoted context omitted.

We'd be happy to provide access to MI300X at TensorWave so you can validate our results! Just shoot us an email or fill out the form on our website

If you're able to advertise available GPU compute in some public forums then it's enough to tell us about the demand of MI300X in cloud ...

You're joking/trolling right? There are literally 10's of thousands of H100s available on gpulist right now, does that mean there's no cloud demand for Nvidia gpus? (I notice from your comment history that you seem to be some sort of bizarre NVDA stan account, but come on, be serious)

Re: AMD's MI300X Outperforms Nvidia's H100 for LLM Inference

#268

Earlier quoted context omitted.

I think there's also a very high prospect of virtual worlds with virtual people (SFW or otherwise) becoming popular, rendered with Apple/META goggles...that could require insane amounts of compute. And this is just one possibility. Relatively cheap multimodal smart glasses you wear when out and around that offload compute to the cloud are another. Nvidia could just as easily triple in short order as get cut in half f…

You mean like Omniverse from Nvidia where you can simulate entire factories or as Nvidia does data centers before they are built? Or how about building the world virtually? https://www.nvidia.com/en-us/high-performance-computing/eart...

Ya, I saw those demos (I own the stock), incredible. But I'm thinking things more like just a single VR friend who has memory and (optionally) prior knowledge of your background. I think these could be an absolute blessing for a lot of people who have lots of time on their hands but no one to talk to.

Or, leaning more towards your examples, a Grand Theft Auto style environment containing millions of them, except life like.

Re: AMD's MI300X Outperforms Nvidia's H100 for LLM Inference

#269
post #76

Given that a lot of projects are written or optimised for CUDA, would it require an industry shift if AMD were to become a competitive source of GPUs for AI training?

The model code is comparatively tiny compared to pytorch or CUDA itself. Translating models from CUDA/C could be laborious but not a barrier. Making AMD work effortlessly with pytorch et al should make the switch transparent.

AMD supports PyTorch out of the box , these comments make me feel no one has tried or even working on this stuff

Re: AMD's MI300X Outperforms Nvidia's H100 for LLM Inference

#270
post #142

Earlier quoted context omitted.

And yet standards of living in Europe are comparable to those in the US, and preferable at the median. Our attention is captured by speculative valuations of unicorns, and yet people actually need real stuff made, drugs developed and made etc. Europe does perfectly well in many non winner takes all sectors where English language and network effects are less relevant. The political instability created by the US neolib…

Europe is much much poorer than the US, actually. https://pbs.twimg.com/media/F3PGpsrWEAEiplB?format=jpg&name=...

These are mean figures, not median.
Post reply on HN