Live data from Hacker News

Why CUDA translation wont unlock AMD

eliovp.com

41–50 of 85 posts

Re: Why CUDA translation wont unlock AMD

#41

[flagged]

No. This is far beyond the capabilities of current AI, and will remain so for the foreseeable future. You could let your model of choice churn on this for months, and you will not get anywhere. It will be able to reach a somewhat working solution quickly, but it will soon reach a point where for every issue it fixes, it introduces one or more issues or regressions. LLMs are simply not capable of scaffolding complexity like a human, and lack the clarity and rigorousness of thought required to execute an *extremely* ambitious project like performant CUDA to ROCm translation.

Re: Why CUDA translation wont unlock AMD

#42

[flagged]

No. This is far beyond the capabilities of current AI, and will remain so for the foreseeable future. You could let your model of choice churn on this for months, and you will not get anywhere. It will be able to reach a somewhat working solution quickly, but it will soon reach a point where for every issue it fixes, it introduces one or more issues or regressions. LLMs are simply not capable of scaffolding complexit…

[flagged]

Re: Why CUDA translation wont unlock AMD

#43
post #9

Let’s just say what it is: devs are too constrained to jump ship right now. It’s a massive land grab and you are not going to spend time tinkering with CUDA alternatives when even a six-month delay can basically kill your company/organization. Google and Apple are two companies with enough resources to do it. Google isn’t because they’re keeping it proprietary to their cloud. Apple still have their heads stuck in san…

Google has their own TPUs so they don’t have any vendor lock-in issues at all. OpenAI OTOH is big enough that the vendor lock-in is actually hurting them, and them making that massive deal with AMD may finally push the needle for AMD and improve things in the ecosystem to make AMD a smooth experience.

Google’s TPU’s are not powering Gemini or whatever X equivalent LLM you want to compare to.

Re: Why CUDA translation wont unlock AMD

#44
post #43

Earlier quoted context omitted.

Google has their own TPUs so they don’t have any vendor lock-in issues at all. OpenAI OTOH is big enough that the vendor lock-in is actually hurting them, and them making that massive deal with AMD may finally push the needle for AMD and improve things in the ecosystem to make AMD a smooth experience.

Google’s TPU’s are not powering Gemini or whatever X equivalent LLM you want to compare to.

What is powering Gemini?

Re: Why CUDA translation wont unlock AMD

#45
post #43

Earlier quoted context omitted.

Google has their own TPUs so they don’t have any vendor lock-in issues at all. OpenAI OTOH is big enough that the vendor lock-in is actually hurting them, and them making that massive deal with AMD may finally push the needle for AMD and improve things in the ecosystem to make AMD a smooth experience.

Google’s TPU’s are not powering Gemini or whatever X equivalent LLM you want to compare to.

This isn't true. Gemini is trained and run almost entirely on TPUs. Anthropic also uses TPUs for inference, see, e.g., https://www.anthropic.com/news/expanding-our-use-of-google-c... and https://www.anthropic.com/engineering/a-postmortem-of-three-.... OpenAI also uses TPUs for inference at least in some measure: https://x.com/amir/status/1938692182787137738?t=9QNb0hfaQShW....

Re: Why CUDA translation wont unlock AMD

#46
post #43

Earlier quoted context omitted.

Google has their own TPUs so they don’t have any vendor lock-in issues at all. OpenAI OTOH is big enough that the vendor lock-in is actually hurting them, and them making that massive deal with AMD may finally push the needle for AMD and improve things in the ecosystem to make AMD a smooth experience.

Google’s TPU’s are not powering Gemini or whatever X equivalent LLM you want to compare to.

I can assure you that most internal ML teams are using TPUs both for training and inference, they are just so much easier to get. Whatever GPUs exist are either reserved for Google Cloud customers, or loaned temporarily to researchers who want to publish easily externally reproducible results.

Re: Why CUDA translation wont unlock AMD

#47
post #43

Earlier quoted context omitted.

Google has their own TPUs so they don’t have any vendor lock-in issues at all. OpenAI OTOH is big enough that the vendor lock-in is actually hurting them, and them making that massive deal with AMD may finally push the needle for AMD and improve things in the ecosystem to make AMD a smooth experience.

Google’s TPU’s are not powering Gemini or whatever X equivalent LLM you want to compare to.

They are, even Apple famously uses Google Cloud for their cloud based AI stuff solely because of Apple not wanting to buy NVidia.

Google Cloud does have a lot of NVidia, but that’s for their regular cloud customers, not internal stuff.

Re: Why CUDA translation wont unlock AMD

#48

https://geohot.github.io//blog/jekyll/update/2025/03/08/AMD-... https://tinygrad.org/ is the only viable alternative to CUDA that I have seen popup in the past few years.

Both Mojo and ThunderKittens/HipKittens are viable on AMD.

Mojo runs faster on nvidia hardware than CUDA in some cases.

https://x.com/clattner_llvm/status/1982196673771139466?s=61

Re: Why CUDA translation wont unlock AMD

#49

Earlier quoted context omitted.

No. This is far beyond the capabilities of current AI, and will remain so for the foreseeable future. You could let your model of choice churn on this for months, and you will not get anywhere. It will be able to reach a somewhat working solution quickly, but it will soon reach a point where for every issue it fixes, it introduces one or more issues or regressions. LLMs are simply not capable of scaffolding complexit…

[flagged]

Well that's your problem. Here's a tip: just because someone says something doesn't mean you have to listen to them
Post reply on HN