Earlier quoted context omitted.
The model code is comparatively tiny compared to pytorch or CUDA itself. Translating models from CUDA/C could be laborious but not a barrier. Making AMD work effortlessly with pytorch et al should make the switch transparent.
These kinds of comments make me think few people have actually tried. My experience has been 1 work day of getting things set up to work the same as before for training and testing (PyTorch).
AMD's MI300X Outperforms Nvidia's H100 for LLM Inference
221–230 of 273 posts
Re: AMD's MI300X Outperforms Nvidia's H100 for LLM Inference
#222Earlier quoted context omitted.
I love AMD, but my Nvidia stock position currently is much higher than AMD.
AMD is at a much higher PE ratio. Is the market expecting AMD to up its game in the GPU sector? Or is the market expecting a pullback in GPU demand due to possibility for non-GPU AI solutions becoming the frontier or for AI investment to slow down?
Re: AMD's MI300X Outperforms Nvidia's H100 for LLM Inference
#223Earlier quoted context omitted.
Maybe I'm a naive fanboy, but I would put my money on Apple catching Nvidia before AMD or Intel.
Apple doesn't have any hardware SIMD technology that I'm aware of. At best, Apple has Metal API which iOS video games use. I guess there's a level of SIMD-compute expertise here, but it'd take a lot of investment to turn that into a full scale GPU that tangos with Supercomputers. Software is a bit piece of the puzzle for sure, but Metal isn't ready for prime time. I'd say Apple is ahead of Intel (Intel keeps wasting…
Re: AMD's MI300X Outperforms Nvidia's H100 for LLM Inference
#224Re: AMD's MI300X Outperforms Nvidia's H100 for LLM Inference
#225Earlier quoted context omitted.
the market and the selling price also includes sales strategies, penetrating a sector dominated by a strong player with somewhat "smart" sales strategies *1 and with a growing but certainly less mature product ( expecially software ), it requires suitable pricing and allocation strategies 1. https://www.techspot.com/news/102056-nvidia-allegedly-punish...
This stuff is the actual reason nvidia is under antitrust investigation. boo boo, a GTX 670 that cost you $399 in 2012 now costs $599 - grow up, do the inflation calculation, and realize you’re being a child. gamers get the best deal on bulk silicon on the planet, R&D subsidized by enterprise, fantastic blue-sky research that takes years for competitors to (not even) match, and it’s still never enough. ”Gamers” have…
Re: AMD's MI300X Outperforms Nvidia's H100 for LLM Inference
#226"TensorWave is a cloud provider specializing in AI workloads. Their platform leverages AMD’s Instinct™ MI300X accelerators, designed to deliver high performance for generative AI workloads and HPC applications." I suggest taking the report with a grain of salt.
The salt is in the plain sight. The do the standard AMD comparison: 8x AMD MI300X (192GB, 750W) GPU 8x H100 SXM5 (80GB, 700W) GPU The fair comparison would be against 8x H100 NVL (188GB, Price tells a story. If AMD performance would be in par with Nvidia they would not sell their cards for 1/4 price.
I don't think it should be ignored, especially when the power consumption is similar.
Re: AMD's MI300X Outperforms Nvidia's H100 for LLM Inference
#227Earlier quoted context omitted.
They can't make them fast enough to satisfy hype-driven demand; they are not scaling like a tech company, but their market cap is being inflated like one.
What is the meaning of "tech" in this context? "Software"? I mean, if Nvidia isn't a technology company...
Re: AMD's MI300X Outperforms Nvidia's H100 for LLM Inference
#228Earlier quoted context omitted.
> Google is not just ads. Bad example given how aggressively they terminate products which don't generate the same revenue as ads. > Apple is not. Best example, they have done a fantastic job of being both a tech company and pseudo-fashion company. > Amazon is not. They don't make anything (at least nothing people want to buy) and have ad revenue as an increase slice of their pie. > Tesla is not. Even bigger hype/spe…
Amazon doesn’t make anything anyone wants to buy? Not to be snarky, but if AWS counts as “nothing” I’d sure like a slice of nothing please.
Re: AMD's MI300X Outperforms Nvidia's H100 for LLM Inference
#229Earlier quoted context omitted.
It's more how little the Frankfurt stock Exchange is worth. And European devs keep wondering why our wages are lower than in the US for the same work. That's why.
The DAX is only 40 companies, most of which make real products rather than advertising mechanisms. Making real physical things just doesn't scale, and never will. While I would enjoy a US tech salary, I'm not sure we want a world where all manufacturing is set aside to focus on the attention economy. Nvidia value deserves to be much higher than any company on the DAX (maybe all of them together, as it currently is) -…
Nvidia sells physical things, and they are bigger than 40 companies because the companies are selling physical things?
I am not arguing hardware scales better than software but this is a strange argument in this context.
Re: AMD's MI300X Outperforms Nvidia's H100 for LLM Inference
#230Earlier quoted context omitted.
> but I think it's one that's mostly about Nvidia's platform dominance and profit margins more Profit margins and dominance are result from performance, not the other way around. It does not matter if Nvidia tools are better when you deploy large number of chips for inference and it does more flops per watt or second. It's seller market and if AMD can't ask high price, their chip do not perform. ---- Question: People…
> People here seem to think that Nvidia has absolutely no advantage in their microarchitecture design skills. It's all in software or monopoly. That's an extrapolation. Microarchitecture design skills are not theoretical numbers you manage to put on a spec sheet. You cannot decouple the software driving the hardware - that's not a trivial problem.
not only can you measure this, not only do they measure this, but it's literally the first component of the Rayleigh resolution equation and everyone is constantly optimizing for it all the time.
https://youtu.be/HxyM2Chu9Vc?t=196
https://www.lithoguru.com/scientist/CHE323/Lecture48.pdf
in the abstract, why does it surprise you that the semiconductor industry would have a way to quantify that?
like, realize that NVIDIA being on a tear with their design has specifically coincided with the point in time when they decided to go all-in on AI (2014-2015 era). Maxwell was the first architecture that showed what a stripped-down architecture could do with neural nets, and it is pretty clear that NVIDIA has been working on this ML-assisted computational lithography and computational design stuff for a while. Since then, I would say - but they've been public about it for several years now (and might be longer, I'd have to look back).
https://www.newyorker.com/magazine/2023/12/04/how-jensen-hua...
https://www.youtube.com/watch?v=JXb1n0OrdeI&t=1383s
Since that "mid 2010s" moment, it's been Pascal vs Vega, Turing (significant redesign and explicit focus on AI/tensor) vs RDNA1 (significant focus on crashing to desktop), Ampere vs RDNA2, etc. Since then, NVIDIA has almost continuously done more with less: beaten custom advanced tech like HBM with commodity products and small evolutions thereupon (like GDDR5X/6X), matched or beaten the efficiency of extremely expensive TSMC nodes with junk samsung crap they got for a song, etc. Quantitatively by any metric they have done much better than AMD. Like Vega is your example of AMD design? Or RDNA1, the architecture that never quite ran stable? RDNA3, the architecture that still doesn't idle right, and whose MCM still uses so much silicon it raises costs instead of lowering them? Literally the sole generation that's not been a total disaster from The Competition has been RDNA2, so yeah, solid wins and iteration is all it takes to say they are doing quantitatively better, especially considering NVIDIA was overcoming a node disadvantage for most of that. They were focused on bringing costs down, and frankly they were so successful despite that that AMD kinda gave up on trying to outprice them.
Contrast to the POSCAP/MLCC problem in 2020: despite a lot of hype from tech media that it was gonna be a huge scandal/cost performance, NVIDIA patched it dead in a week with basically no perf cost etc. Gosh do you think they might have done some GPGPU accelerated simulations to help them figure that out so quickly, how the chip was going to boost and what the transient surges were going to be etc?
literally they do have better design skills, and some of it is their systems thinking, and some of it is their engineers (they pay better/have better QOL and working conditions, and get the cream of the crop), and some of it is their better design+computational lithography techniques that they have been dogfooding for 3-4 generations now.
people don't get it: startup mentality, founder-led, with a $3t market cap. Jensen is built different. Why wouldn’t they have been using this stuff internally? That’s an extremely Jensen move.