Live data from Hacker News

CUDA Moat Still Alive

semianalysis.com

81–90 of 176 posts

Re: CUDA Moat Still Alive

#81

Earlier quoted context omitted.

I don't have the inside baseball but I have seen those weird as hell interviews with Lisa Su where she gets asked point blank about the software problems and instead of "working on it, stay tuned" -- an answer that costs nothing to give -- she deflects into "performance is what matters," which is the kind of denial that rhymes exactly with the problems they are having. No, the horsepower of your F1 racecar doesn't ma…

> she deflects into "performance is what matters," which is the kind of denial that rhymes exactly with the problems they are having. It's not a deflection, but a straightforward description of AMDs current top-down market strategy of partnering with big players instead of doubling down to have a great OOBE for consumers & others who don't order GPUs by the pallet. It's an honest reflection if their current core comp…

>Hyperscalers are also more self-sufficient at software: they have entire teams working on PyTorch, Jax, and writing kernels.

None of this matters because AMD drivers are broken. No one is asking AMD to write a PyTorch backend. The idea that AMD will have twice the silicon performance than nvidia to make up the performance loss for bad software is a pipedream.

Re: CUDA Moat Still Alive

#82

AMD could spend their market cap in one year to get this done in three and it would be a coup for the shareholders. They could hire all of the best NVIDIA engineers at double their current comp, crush the next TSMC node on Apple levels, and just do it and if it got them a quarter of NVDA’s cap it would be a bargain. They don’t fucking want to! Believing this is anything like a market is fucking religion.

Try to make sense...

They can spend their market cap by either:

1: issuing new shares worth their market cap, diluting existing shareholders to 50%.

2: Or borrow their market cap and pay interest by decreasing profits. "AMD operating margin for the quarter ending September 30, 2024 was 5.64%" so profits would be extremely impacted by interest repayments.

Either way your suggestion would be unlikely to be supported by shareholders.

> crush the next TSMC node on Apple levels

I would guess Apple is indirectly paying for the hardware (to avoid repatriating profits) or guaranteeing usage to get to the front of the line at TSMC. Good luck AMD competing with Apple: there's a reason AMD sold GlobalFoundries and there's a reason Intel is now struggling with their foundry costs.

And it comes across as condescending to assume you know better than a successful company.

Re: CUDA Moat Still Alive

#83
post #82

AMD could spend their market cap in one year to get this done in three and it would be a coup for the shareholders. They could hire all of the best NVIDIA engineers at double their current comp, crush the next TSMC node on Apple levels, and just do it and if it got them a quarter of NVDA’s cap it would be a bargain. They don’t fucking want to! Believing this is anything like a market is fucking religion.

Try to make sense... They can spend their market cap by either: 1: issuing new shares worth their market cap, diluting existing shareholders to 50%. 2: Or borrow their market cap and pay interest by decreasing profits. "AMD operating margin for the quarter ending September 30, 2024 was 5.64%" so profits would be extremely impacted by interest repayments. Either way your suggestion would be unlikely to be supported by…

When 4 trillion dollars are at stake, the financing is available or it fucking better be.

What in God’s name do we pay these structured finance, bond-issue assholes 15% of GDP for if not to finance a sure thing like that?

It sure as hell ain’t for their taste in Charvet and Hermes ties, because the ones they pick look like shit.

Re: CUDA Moat Still Alive

#84

Earlier quoted context omitted.

I don't have the inside baseball but I have seen those weird as hell interviews with Lisa Su where she gets asked point blank about the software problems and instead of "working on it, stay tuned" -- an answer that costs nothing to give -- she deflects into "performance is what matters," which is the kind of denial that rhymes exactly with the problems they are having. No, the horsepower of your F1 racecar doesn't ma…

> she deflects into "performance is what matters," which is the kind of denial that rhymes exactly with the problems they are having. It's not a deflection, but a straightforward description of AMDs current top-down market strategy of partnering with big players instead of doubling down to have a great OOBE for consumers & others who don't order GPUs by the pallet. It's an honest reflection if their current core comp…

Engineers at hyperscalers are struggling through all the bugs too. It's coming at notable opportunity cost for them, at a time when they also want an end to the monopoly. Do they buy AMD and wade through bug after bug, regression after regression, or do they shell out slightly more money for Nvidia GPUs and have it "just work".

AMD has to get on top of their software quality issues if they're ever going to succeed in this segment, or they need to be producing chips so much faster than Nvidia that it's worth the extra time investment and pain.

Re: CUDA Moat Still Alive

#85
post #56
post #54

> AMD is attempting to vertically integrate next year with their upcoming Pollara 400G NIC, which supports Ultra Ethernet, hopefully making AMD competitive with Nvidia. Infiniband is an industry standard. It is weird to see the industry invent yet another standard to do effectively the same thing just because Nvidia is using it. This “Nvidia does things this way so let’s do it differently” mentality is hurting AMD: *…

You're painting this like AMD is off to play in their own sandbox when it's more like the entire industry is trying to develop an alternative to Infiniband. Ultra Ethernet is a joint project between dozens of companies organized under the Linux Foundation. https://www.phoronix.com/news/Ultra-Ethernet-Consortium >> The Linux Foundation has established the Ultra Ethernet Consortium "UED" as an industry-wide effort foun…

I wrote:

> It is weird to see the industry invent yet another standard to do effectively the same thing just because Nvidia is using it.

This is a misstep for all involved, AMD included. Even if AMD is following everyone else by jumping off a bridge, AMD is still jumping too.

Re: CUDA Moat Still Alive

#86

Earlier quoted context omitted.

That's the excuse used by every big company shitting out software so broken that it needs intensive professional babysitting. I've been on both sides of this shitshow, I've even said those lines before! But I've also been in the trenches making the broken shit work and I know that it's fundamentally an excuse. There's a reason why people pay 80% margin to Nvidia and there's a reason why AMD is worth less than the rou…

What exactly are they in denial about? They are aware that software is not a strength of theirs, so they partner with those who are great at it. Would you say AMD is "shitting the bed" by not building it's own consoles too? You know AMD could build a kick-ass console since they are doing the heavy-lifting for the Playstation, and the XBox[1] , but AMD knows as much as anybody that they don't have the skills to wrangl…

There is probably one employee - either a direct report of Su's or maybe one of her grandchildren in the org chart - who needs to "get it". If they replaced that one manager with someone who sees graphics cards as a tool to accelerate linear algebra then AMD would be participating more effectively in a multi-trillion dollar market. They are so breathtakingly close to the minimum standards of competence on this one. We know from the specs that the cards they produce should be able to perform.

This is a case-specific example of failure, it doesn't generalise very well to other markets. AMD is really well positioned for this very specific opportunity of historic proportions and the only thing holding them back is a somewhat continuous stream of unforced failures when writing a high quality compute driver. It seems to be pretty close to one single team of people holding the company back although organisational issues tend to stem from a level or two higher than the team. This could be the most visible case of value destruction by a public company we'll see in our lifetimes.

Optimistically speaking maybe they've already found and sacked the individual responsible and we're just waiting for improvement. I'm buying Nvidia until that proves to be so.

Re: CUDA Moat Still Alive

#87

Earlier quoted context omitted.

> she deflects into "performance is what matters," which is the kind of denial that rhymes exactly with the problems they are having. It's not a deflection, but a straightforward description of AMDs current top-down market strategy of partnering with big players instead of doubling down to have a great OOBE for consumers & others who don't order GPUs by the pallet. It's an honest reflection if their current core comp…

> Hyperscalers are also more self-sufficient at software: they have entire teams working on PyTorch, Jax, and writing kernels. None of this matters because AMD drivers are broken. No one is asking AMD to write a PyTorch backend. The idea that AMD will have twice the silicon performance than nvidia to make up the performance loss for bad software is a pipedream.

> None of this matters because AMD drivers are broken

Do you honestly think the MI300 has show-stopper driver bugs, or that Meta/Amazon doesn't have a direct line to AMD engineers?

Re: CUDA Moat Still Alive

#88
post #15

> CUDA Moat Still Alive Wrong conclusion. AMD is slower than NVidia, but not _that_ much slower. They are actually pretty cost-competitive. The just need to do some improvements, and they'll be a very viable competitor.

All the cloud providers list MI300x as more expensive than H100. So if you compare performance/cost it is even worse.

Re: CUDA Moat Still Alive

#89
post #32

That MatMul performance is fairly shocking. To be that much below theoretical maximum on what should be a fairly low overhead operation. I would at least hope that they know where the speed is going, but the issue of torch.matmul and F.Linear using different libraries with different performance suggests that they don't even know which code they are running, let alone where the slow bits in that code are.

Low overhead in what sense? matmul is kinda complicated and there are varying, complex state-of-the-art algorithms for it, no? And then if you know things about the matrices in advance you can start optimizing for that, which adds another layer of complexity.

There are, but everyone uses variations of the same O(n^3) algorithm taught in college introduction to linear algebra classes because it is numerically stable and can be made extremely fast through tweaks that give spatial locality and good cache characteristics. Meanwhile the asymptomatically faster algorithms have such large constants in their big O notation that they are not worth using. FFT based matrix multiplication, which is O((n^2)log(n)), also has numerical instability on top of running slower.

Re: CUDA Moat Still Alive

#90
post #76
post #66

Earlier quoted context omitted.

By overhead I'm talking about the things that have to be done supplementary to the algorithm. While there are complex state-of-the-art algorithms, those algorithms exist for everyone. The overhead is the bit that had to be done to make the algorithm work. For instance for sorting a list of strings the algorithm might be quick sort. The overhead would be in the efficiency of your string compare. For matmul I'm not sur…

> For matmul I'm not sure what your overhead is beyond moving memory, multiplying, and adding. A platform touting a memory bandwidth and raw compute advantage should have that covered. Where is the performance being lost? The use of the word 'algorithm' is incorrect. Look... I do this sort of work for a living. There has been no useful significant change to matmul algorithms . What has changed is the matmul process .…

To be fair, GEMV is memory bandwidth bound and that is what token generation in transformers uses. GEMM is the compute bound one, provided you do not shoehorn GEMV into it. That special case is memory bandwidth bound.
Post reply on HN