Live data from Hacker News

AMD's MI300X Outperforms Nvidia's H100 for LLM Inference

blog.tensorwave.com

271–273 of 273 posts

Re: AMD's MI300X Outperforms Nvidia's H100 for LLM Inference

#271
post #198

Earlier quoted context omitted.

You can't get any more obvious than that, but that wasn't my point to say that higher GDP doesn't make you richer than a low GDP, but to say GDP/capita as a number alone is not a measure of wealth, income or prosperity between countries, even in the EU. For example Ireland has by a long margin the highest GDP/capita in the whole EU, and it would make you think the average Irish worker earns more that any other worker…

> For example Ireland has by a long margin the highest GDP/capita in the whole EU No, Luxembourg does.

But in both cases it's because of weird corporate tax loopholes rather than the underlying productivity or utility to the median individual.

Re: AMD's MI300X Outperforms Nvidia's H100 for LLM Inference

#272

Earlier quoted context omitted.

I've used ChatGPT on Azure. It sucks on so many levels, everything about it was clearly enforced by some bean counters who see X dollars for Y flops with zero regard for developers. So choosing AMD here would be about par for the course. There is a reason why everyone at the top is racing to buy Nvidia cards and pay the premium.

"Everyone" at top is also developing their own chips for inference and providing APIs for customers to not worry about using CUDA. It looks like the price to performance of inference tasks gives providers a big incentive to move away from Nvidia.

There are only like 3 AI building companies who have the tech capability and resources to afford that and 2 of them don't even offer their chips to others or have gone back to Nvidia. The rest is manufacturers desperately trying to get a piece of the pie.

Re: AMD's MI300X Outperforms Nvidia's H100 for LLM Inference

#273

Earlier quoted context omitted.

Even then it's an insanely bold assumption that a company other than Nvidia could build a better framework than CUDA for compiling PTX. Especially since CUDA is already so much C-like. I've never seen anyone go deeper than that outside of academia.

How many customers/consumers will care about services are be built with CUDA? If they need a ChatBot that uses a model with same accuracy and performance as on non-CUDA hardware, would they still want CUDA based hardware?

Who is going to build the architecture and compile the device specific kernels? You have to pay those people as well and you can save tons of money and time if you do it with cuda.
Post reply on HN