"Furthermore, AMD OLMo models were also able to run inference on AMD Ryzen™ AI PCs that are equipped with Neural Processing Units (NPUs). Developers can easily run Generative AI models locally by utilizing the AMD Ryzen™ AI Software." Hope these AI PCs will run also something better than 1B model. What is it useful for ? Spellcheck ?
AMD Open-Source 1B OLMo Language Models
11–20 of 35 posts
Re: AMD Open-Source 1B OLMo Language Models
#12Earlier quoted context omitted.
The point is that AMD is doing the legwork to ensure that AI models can run on their chips. While they could settle for inference workloads (port llama to AMD). It is unlikely that many teams will widely adopt their silicon unless they can be used in the end-end ML stack. Many pure OSS efforts have tried and failed to make AMD work for this use case. As a chip maker - they will also have some undersold, QA, or otherw…
It's amazing how NVidia became worth $3T simply because they have better drivers and CUDA. AMD has great hardware, but they never could be assed to do anything about their software.
Re: AMD Open-Source 1B OLMo Language Models
#13Earlier quoted context omitted.
It's amazing how NVidia became worth $3T simply because they have better drivers and CUDA. AMD has great hardware, but they never could be assed to do anything about their software.
It's not really. Anyone who's ever done any low-level assembly coding on modern chips knows that it is already a herculean engineering effort. The idea that your customers, who are experts in machine learning models (like transformers, activation functions, etc) are going to feel comfortable with memory hierarchies, synchronization, floating point precision, etc is just crazy.
AMD did approximately nothing with ROCm.
Investing $10-20m of developer time into making ROCm work reliably easily would have paid for itself 100x.
Re: AMD Open-Source 1B OLMo Language Models
#14Earlier quoted context omitted.
It's not really. Anyone who's ever done any low-level assembly coding on modern chips knows that it is already a herculean engineering effort. The idea that your customers, who are experts in machine learning models (like transformers, activation functions, etc) are going to feel comfortable with memory hierarchies, synchronization, floating point precision, etc is just crazy.
Yes, that's what I mean. NVidia provided easy to use tooling (CUDA), and made sure it JustWorks everywhere. AMD did approximately nothing with ROCm. Investing $10-20m of developer time into making ROCm work reliably easily would have paid for itself 100x.
Re: AMD Open-Source 1B OLMo Language Models
#15Re: AMD Open-Source 1B OLMo Language Models
#16Earlier quoted context omitted.
It's not really. Anyone who's ever done any low-level assembly coding on modern chips knows that it is already a herculean engineering effort. The idea that your customers, who are experts in machine learning models (like transformers, activation functions, etc) are going to feel comfortable with memory hierarchies, synchronization, floating point precision, etc is just crazy.
Yes, that's what I mean. NVidia provided easy to use tooling (CUDA), and made sure it JustWorks everywhere. AMD did approximately nothing with ROCm. Investing $10-20m of developer time into making ROCm work reliably easily would have paid for itself 100x.
I love when outsiders throw around random-ass takes like this. Just curious: how'd you come up with this number? Is it backed by literally any thought/data/roadmap?
Let's do some rough back of the envelope calculations: 20MM is 100 engineers working for 1 year. Or maybe it's 5 years of work for 20 engineers? Which one of those perspectives (if any!) sounds to you like a reasonable assessment of the gap between AMD and NVIDIA?
A quick reminder before you answer: whatever you think is actually involved in improving ROCm, unless you work on ROCm, you're almost certainly not considering an entire iceberg of complexity (runtime/driver/firmware).
Let's put it another way: forget AMD investing, I'll invest in you since you're so confident. I'll give you 20MM as a high-interest, non-dischargeable loan (say 8%) and all the runtime/driver/firmware source for AMDGPU. Up for it? All you have to do is improve ROCm such that it's competitive with CUDA and you can take home a huge slice of the TAM and you'll be rich. Easy right?
Cutting to the chase: you're off by at least two orders of magnitude on your goofy estimate; the real numbers are probably closer to 200MM invested every year for 10 years. And you still wouldn't be caught up because in those 10 years NVIDIA wasn't sitting on its laurels just waiting for you to catch up!
Re: AMD Open-Source 1B OLMo Language Models
#17Earlier quoted context omitted.
Yes, that's what I mean. NVidia provided easy to use tooling (CUDA), and made sure it JustWorks everywhere. AMD did approximately nothing with ROCm. Investing $10-20m of developer time into making ROCm work reliably easily would have paid for itself 100x.
> Investing $10-20m of developer time into making ROCm work reliably easily would have paid for itself 100x. I love when outsiders throw around random-ass takes like this. Just curious: how'd you come up with this number? Is it backed by literally any thought/data/roadmap? Let's do some rough back of the envelope calculations: 20MM is 100 engineers working for 1 year. Or maybe it's 5 years of work for 20 engineers? W…
Re: AMD Open-Source 1B OLMo Language Models
#18Earlier quoted context omitted.
> Investing $10-20m of developer time into making ROCm work reliably easily would have paid for itself 100x. I love when outsiders throw around random-ass takes like this. Just curious: how'd you come up with this number? Is it backed by literally any thought/data/roadmap? Let's do some rough back of the envelope calculations: 20MM is 100 engineers working for 1 year. Or maybe it's 5 years of work for 20 engineers? W…
200 mm/year gets you roughly 1000 engineers at 200k salary. Is that not enough to make rocm experience equal to cuda?
I see it less as an engineering problem and more as a market problem. AMD stuff has existed, it’s the market that doesn’t see a point in it, and at this point, even feature parity or CUDA compatibility for that matter won’t make a huge dent. People will just keep using what they know and are recommended.
It’s more amazing to me that NVDA is so intensely inflated by this LLM hype wave. I find it genuinely scary to think about what’s going to happen when 95+% of AI slopware startups fold. Nvidia won’t be the only company financially impacted. Our entire economy runs on fads.
Re: AMD Open-Source 1B OLMo Language Models
#19Earlier quoted context omitted.
Yes, that's what I mean. NVidia provided easy to use tooling (CUDA), and made sure it JustWorks everywhere. AMD did approximately nothing with ROCm. Investing $10-20m of developer time into making ROCm work reliably easily would have paid for itself 100x.
> Investing $10-20m of developer time into making ROCm work reliably easily would have paid for itself 100x. I love when outsiders throw around random-ass takes like this. Just curious: how'd you come up with this number? Is it backed by literally any thought/data/roadmap? Let's do some rough back of the envelope calculations: 20MM is 100 engineers working for 1 year. Or maybe it's 5 years of work for 20 engineers? W…
Re: AMD Open-Source 1B OLMo Language Models
#20Earlier quoted context omitted.
> Investing $10-20m of developer time into making ROCm work reliably easily would have paid for itself 100x. I love when outsiders throw around random-ass takes like this. Just curious: how'd you come up with this number? Is it backed by literally any thought/data/roadmap? Let's do some rough back of the envelope calculations: 20MM is 100 engineers working for 1 year. Or maybe it's 5 years of work for 20 engineers? W…
I appreciate this comment keeping us in line.