Live data from Hacker News

Attention Decode on AMD MI450 GPUs: A Gluon Kernel Optimization Guide

rocm.blogs.amd.com

11–12 of 12 posts

Re: Attention Decode on AMD MI450 GPUs: A Gluon Kernel Optimization Guide

#11
post #9
post #8

Earlier quoted context omitted.

This thing has 50% more memory than GTX PRO 6000, which fetch 18k-25k+ each right now on newegg, and generally sold out. If nvdia can still sell out 96GB PCI cards (on a 2 year old architecture) at those prices, how is an AMD equivalent with better specs dead on arrival?

Market takes time to catch up. When released Qwen 3.6 had a lot of sense and it did fit, but frontier since went too far ahead. Now you'd want DS V4 Flash from yesterday, or you are too far behind. Unless MI350P will be much cheaper than 6000 Pro, which is unlikely given its VRAM, it will lose to it, because you'd need 2 of either.

Don't see how it could be cheaper with HBM than the RTX 6000 Blackwell Server Pro but in these times of relative shortage any additionnal supply should be slurped ?

I was remarking on the MI350P because I've had a hard time procuring "small" CDNAx systems (for e.g. development, experiments and lower-profile servers) and OAM seemed very niche (not if you're aiming for density and training/inference...).

I hope the MI350P fills a lower part of the spectrum and I can start massively porting CUDA stuff or at least work on HIP and ROCm and what I need to make most or some of our CUDA stuff run on AMD HW, then how make it run fast.

Re: Attention Decode on AMD MI450 GPUs: A Gluon Kernel Optimization Guide

#12

As much as I enjoy these articles and for AMD to write more light technical articles, it really feels constrained, even strained, to be unable to cite the equivalent terms from the precursor here (NVIDIA). Another batch of jargon for very similar architectures and programming models... HIP and ROCm have actually made amazing strides in making CUDA developers' porting work easy, and I know playing catchup to a (monopo…

Most Triton terminology comes from the AMD engineers working on it, rather than Nvidia's names.
Post reply on HN