CUDA Moat Still Alive
31–40 of 176 posts
Re: CUDA Moat Still Alive
#32I would at least hope that they know where the speed is going, but the issue of torch.matmul and F.Linear using different libraries with different performance suggests that they don't even know which code they are running, let alone where the slow bits in that code are.
Re: CUDA Moat Still Alive
#33[flagged]
I don't think many people can keep track of who's running Intel let alone have hope that with a little work they can deliver reasonable substitutes for NVIDIA's products.
Re: CUDA Moat Still Alive
#34> CUDA Moat Still Alive Wrong conclusion. AMD is slower than NVidia, but not _that_ much slower. They are actually pretty cost-competitive. The just need to do some improvements, and they'll be a very viable competitor.
The amount of effort this team took, literally co-opting AMD engineers, and working for 5 months, to get closer but not yet usable, means they are not even close to usable. What team wanting to do ML training/inference can afford so much down time for zero benefit? How many except a few big ones can get AMD to devote so many resources simply for that team? And, if you’re training a model costing you millions, the las…
Meanwhile, Nvidia hardware is expensive and still is in short supply. AMD might look quite tempting.
Re: CUDA Moat Still Alive
#35> Give AMD Engineers more compute and engineering resources to fix and improve the AMD ecosystem, they have very few internal gpu boxes relative to what Nvidia provides to their engineers. This is real. We’ve found ourselves having to give hardware to engineers at AMD because they’re unable to get allocation of it internally.
Coming up next: "We bought AMD stock on the open market and used it to compensate AMD engineers".
Spend a billion on AMD shares, Spend another Billion on a out-of-house software team to solve the software solution to more than double the share price.
Taking into account that there are players that already own billions in AMD shares, they could probably do that as well. On the other hand perhaps it would be better for them, as major shareholders, to have a word with AMD management.
Re: CUDA Moat Still Alive
#36That MatMul performance is fairly shocking. To be that much below theoretical maximum on what should be a fairly low overhead operation. I would at least hope that they know where the speed is going, but the issue of torch.matmul and F.Linear using different libraries with different performance suggests that they don't even know which code they are running, let alone where the slow bits in that code are.
Re: CUDA Moat Still Alive
#37> Give AMD Engineers more compute and engineering resources to fix and improve the AMD ecosystem, they have very few internal gpu boxes relative to what Nvidia provides to their engineers. This is real. We’ve found ourselves having to give hardware to engineers at AMD because they’re unable to get allocation of it internally.
Re: CUDA Moat Still Alive
#38Earlier quoted context omitted.
The amount of effort this team took, literally co-opting AMD engineers, and working for 5 months, to get closer but not yet usable, means they are not even close to usable. What team wanting to do ML training/inference can afford so much down time for zero benefit? How many except a few big ones can get AMD to devote so many resources simply for that team? And, if you’re training a model costing you millions, the las…
Sure. But this work is done, and can be reused by others. Meanwhile, Nvidia hardware is expensive and still is in short supply. AMD might look quite tempting.
Meanwhile those libs release running CUDA on NVidia’s old and newest releases out of the box.
So no, it cannot be reused by others in production any more than my custom hacked car engine mod can be added by Ford to every car in existence.
Have you done any deep professional production work on any of these stacks? I have, and would never, ever put stuff like the stuff in the article in production. It’s no where near ready for production use.
Re: CUDA Moat Still Alive
#39> Give AMD Engineers more compute and engineering resources to fix and improve the AMD ecosystem, they have very few internal gpu boxes relative to what Nvidia provides to their engineers. This is real. We’ve found ourselves having to give hardware to engineers at AMD because they’re unable to get allocation of it internally.
Sadly common at hardware companies. The most extreme case I've heard of is ASML, who supposedly doesn't keep any machines of their own. They test against "almost-ready" machines right before they go out the door to customers.
Re: CUDA Moat Still Alive
#40> Give AMD Engineers more compute and engineering resources to fix and improve the AMD ecosystem, they have very few internal gpu boxes relative to what Nvidia provides to their engineers. This is real. We’ve found ourselves having to give hardware to engineers at AMD because they’re unable to get allocation of it internally.