Live data from Hacker News

HipKittens: Fast and furious AMD kernels

hazyresearch.stanford.edu

81–90 of 94 posts

Re: HipKittens: Fast and furious AMD kernels

#81
post #14

Earlier quoted context omitted.

AMD pays the bare minimum in software to get a product out the door. The company does not even have working performance testing and regressions routinely get shipped to customers. Benchmarks the executives see are ad hoc and not meaningful. HipKittens is an improvement but AMD does not have the ability to understand or track kernel performance so it'll be ignored. This isn't fixable overnight. Company-wide DevOps and…

> We haven't had a fully funded bonus in the past 4+ years. This is WILD to hear considering how well it appears AMD is executing from the outside.

> considering how well it appears AMD is executing from the outside.

The party line is that the stock price is up because the market expects us to perform well in the future, and we won't get a bonus until we actually perform well.

Re: HipKittens: Fast and furious AMD kernels

#82
post #9

Earlier quoted context omitted.

From the performance comparison table, basically AMD could be NVIDIA right now, but they aren’t because… software? That’s a complete institutional and leadership failure. Ironically, building chips is the actual _hard_ part. The software and the compilers are not trivial but the iteration speed is almost infinite by comparison. It goes to show that some companies just don’t “get” software. Not even AMD!

CUDA was started in 2004. AMD was basically broke until they hit a home run with Ryzen in 2017.

Funnily enough AMD was actually the first with GPGPU... they just floundered and managed to start 3 or more completely new software stacks for it, while CUDA focused not just on keeping one backward compatible one, but also made it work from cheapest NVS card to high end parts.

Re: HipKittens: Fast and furious AMD kernels

#83
post #78

> what is raw assembly? can't understand it? that's the point! Raw assembly vs cooked assembly? Also, I think this attitude wasn’t the most common on CPUs, and people used to write assembly by hand just fine (and sometimes some still do). I think we shouldn’t be afraid of assembly like that. Compilers could write that assembly in the end, just like the do for CPUs!

Yeah, comments like these really make you question the authors' background in optimization. Never mind that AMD actually publishes ISA specs for all of their graphics IPs -- it is not their point that you don't understand it -- what's holding GPU programming back is often that the underlying assembly primitives are not exposed in the high level languages.

I also do wonder what 'raw assembly' is supposed to be. Is it like sushi? Perhaps it is left as future work in the paper for the authors to answer.

Re: HipKittens: Fast and furious AMD kernels

#84
post #45

Earlier quoted context omitted.

As CEO of an AMD NeoCloud for the past 2 years, it is so nice to hear all this and also see the turn around. It is what I bet my business on from the start and I can concur with what George is saying 100%. The out of box experience can be a bit rough around the edges on bleeding edge stuff, but it isn't anything near as bad as it used to be. For example, a month ago nanochat wasn't working well and now it is. The imp…

That’s interesting, I was specifically looking for AMD hardware being offered by neoclouds, they seem to be rare. I like your bet though. The difference between NVDA and AMD has never really existed on a hardware level for decades. AMD has always been on par, and software is software, it will catch up. AMD will be a stock many people will miss because the opportunity has presented itself at the height of AI bubble ta…

Will it catch up or will it forever chase nvidia's tail? I'm betting on the latter unless another AI winter happens. And contrary to anti-generative AI social media talking points, the literature suggests The Red Queen's race is continuing apace IMO.

Nvidia remains undefeated at responding to hardware threats with hardware diving catches to this day. What scenario prevents them from yet another one of their diving catches? I'm genuinely curious as to how one could pull that off. It's like challenging Google in search: even if you deliver better product and some have, the next thing you know Google is doing the same thing or better with deeper pockets.

Re: HipKittens: Fast and furious AMD kernels

#85
post #3

Earlier quoted context omitted.

Infiniband is being replaced with UEC (and it isn't needed for inference). For inference there is no moat and smart players are buying/renting AMD or Google TPUs.

I didn't know you can you buy Google TPUs now?

What you really want to buy is a Ming Mecca chip. Original model came out around 2003, but they've been iterating. These things are bigger than AMD or nvidia silicon, actually even much larger than a gigantic Cerebras wafer, typically 500-900 million USD in price. As you could guess, Ming Mecca is not broadly publicized, historically used for NSA crypto cracking although now adapted to AI and used for data crunching from gathered messages. More recently all those gathered messages have been used for training strategic /tactical intelligence developments to oversee and deploy resources optimally via a cluster of, at least last I heard, 18 Ming Mecca v7 chips

Re: HipKittens: Fast and furious AMD kernels

#86
post #69

With these new developments, are there any implications for getting LLMs running well on consumer AMD chips ? For example, the following laptop which I'm thinking of picking up, has both a strong AMD CPU/IGPU and a RTX 5080. Could we see the AMD side competing with the RTX? I know a dedicated gpu will always be faster though. >HP OMEN MAX 16-ak0003nr 16" Gaming Laptop Computer - Shadow Black Aluminum AMD Ryzen AI 9 H…

You might think that a dGPU is always faster but the limited memory capacity bites you there (unless you go to datacenter dGPUs that cost tens of thousnds). Look at eg https://www.ywian.com/blog/amd-ryzen-ai-max-plus-395-native-... or the various high end Mac results.

So I want this Thinkpad.

https://www.lenovo.com/us/en/p/laptops/thinkpad/thinkpadp/th...?

AMD Ryzen™ AI 9 HX PRO 370 Processor (2.00 GHz up to 5.10 GHz) Operating System Windows 11 Pro 64 Graphic Card Integrated AMD Radeon™ 890M Memory 64 GB DDR5-5600MT/s (SODIMM)(2 x 32 GB)

But I also seriously want to run LLMs. My hunch is a gaming laptop is the best way to do this on the go without spending 5000$ for a Thinkpad with a high end graphics card.

Re: HipKittens: Fast and furious AMD kernels

#87

Earlier quoted context omitted.

That's great I have been eyeing a Strix Halo and was wondering how well smaller models are doing. This is great news from the perspective of running local agents.

I got one of those running whisper yesterday, hopeful the bigger llms will run shortly. You'd need rocm 7 which seems to be much better than 6.4 was.

Is the performance decent? I'm looking at using it with 30b coding models with a local agent framework like goose to see if we can do this locally as developers instead of risking leaking code to the big models.

Re: HipKittens: Fast and furious AMD kernels

#88
post #45

Earlier quoted context omitted.

That’s interesting, I was specifically looking for AMD hardware being offered by neoclouds, they seem to be rare. I like your bet though. The difference between NVDA and AMD has never really existed on a hardware level for decades. AMD has always been on par, and software is software, it will catch up. AMD will be a stock many people will miss because the opportunity has presented itself at the height of AI bubble ta…

Will it catch up or will it forever chase nvidia's tail ? I'm betting on the latter unless another AI winter happens. And contrary to anti-generative AI social media talking points, the literature suggests The Red Queen's race is continuing apace IMO. Nvidia remains undefeated at responding to hardware threats with hardware diving catches to this day. What scenario prevents them from yet another one of their diving c…

Nvidia remains undefeated at responding to hardware threats with hardware diving catches to this day. What scenario prevents them from yet another one of their diving catches?

The fact that they make roughly the same hardware as AMD for the last 2 decades, and even today. There was no diving catch, AMD just ignored what the hardware was capable of and didn't reinforce OpenCL. There was literally no diving catch. For example, just in this thread alone, AMD paid someone to make this shit work on their hardware. Don't bet against what's coming.

Re: HipKittens: Fast and furious AMD kernels

#89
post #5

Ahh, composable-kernel. The highest offender in the list of software that have produced unrecoverable OOMs in my Gentoo system (it’s actually Clang while compiling CK, which uses upwards of 2.5GB per thread).

I was recently reviewing a CK package for Debian. My test build crashed due to OOM using -j32 on a 64GB workstation, so I tried with -j1 to be safe. That completed successfully after 190 hours! I think I may need to reduce the number of architectures it's built for to successfully compile it on the official Debian buildd infrastructure, but my (unverified) understanding is that most of its reverse dependencies only n…

Same, -j32 with 64GB on a 3950x. I use 50% of ZRAM, but it’s still not enough most of the times, so I had to make a config called less-threads that only uses 24, with ZRAM enabled.

I also use OOMD, but I have to work on separating my systemd units better, OOMD has killed my greetd session before, and with that my entire tree of userland processes :D

Re: HipKittens: Fast and furious AMD kernels

#90
post #71
post #5

Ahh, composable-kernel. The highest offender in the list of software that have produced unrecoverable OOMs in my Gentoo system (it’s actually Clang while compiling CK, which uses upwards of 2.5GB per thread).

Spending >10 minutes doing template instantiation for a single kernel for a single ISA is impressive! `device_grouped_conv2d_fwd_xdl_ngchw_gkcyx_ngkhw_f16_instance`, what are you doing to our poor friend clang?

And they say Rust is slow!
Post reply on HN