Earlier quoted context omitted.
Don't underestimate the amount of shit people would be willing to deal with to make stuff work. A capable GPU with 24+ GB would sell if it significantly undercuts Nvidia. Just look at geohot building his tinyboxes with AMD cards.
I would personally love that project but there are already so many versioning issues in the space it would be a nightmare if ROCm randomly broke things all the time.
Micron Kicks Off Production of HBM3E Memory
31–39 of 39 posts
Re: Micron Kicks Off Production of HBM3E Memory
#32Earlier quoted context omitted.
If all you want is VRAM you can get old P40s with 24GB for $175 and it's 144GB for $1050. Then you need a big machine to put six of them in but that doesn't cost $4000. But all of this is kludges. The Radeon RX 7900 XTX has more than twice the memory bandwidth of the M3 Max with much better performance per watt than an array of P40s. What you want is that with more VRAM, not any of this misery.
That checks out in principle, but given that P40 doesn't support NVLink, I wouldn't count too much on using six of them together in a performant manner. But yeah the best option remains an MI300 if you can afford that.
https://www.reddit.com/r/LocalLLaMA/comments/142rm0m/llamacp...
Re: Micron Kicks Off Production of HBM3E Memory
#33Earlier quoted context omitted.
I would personally love that project but there are already so many versioning issues in the space it would be a nightmare if ROCm randomly broke things all the time.
I agree, ROCm seems to be a mess from the outside, but I'm glad people are putting in the effort.
Intel could very easily just put a buttload of VRAM on their existing GPUs to stick it to their competitors and make out like bandits. All they'd have to do is charge a Big markup instead of an Enterprise markup. And Intel has a better history of not making broken libraries.
Re: Micron Kicks Off Production of HBM3E Memory
#34How about higher capacity GDDR6X? 48GB consumer cards (or 96GB pro cards) would sell like hotcakes if AMD/Intel dare to break the artificial VRAM segmentation status quo.
Re: Micron Kicks Off Production of HBM3E Memory
#35How about higher capacity GDDR6X? 48GB consumer cards (or 96GB pro cards) would sell like hotcakes if AMD/Intel dare to break the artificial VRAM segmentation status quo.
I swear I'm getting Deja Vu right now, I coulda sworn I've seen this thread before. There's gonna be a guy commenting that "you don't need it" somebody else saying "but I want it!" And a few trying to figure out the economics of it and whether or not it makes any sense. Personally I'd love to have as much VRAM as possible (and as high a bandwidth as is possible too) to mess around with simulations in- but that's defi…
Re: Micron Kicks Off Production of HBM3E Memory
#36Earlier quoted context omitted.
That machine with 128GB is $5000, with 48GB is still well over $3000, and has as much memory bandwidth as a $400 GPU. At the current spot price, 128GB of GDDR6 is <$400 and 48GB is <$150, implying that they could be paired with any existing <$1000 GPU to produce something significantly faster for dramatically less money. If anyone could be bothered to make one.
I agree with you, that’s how it should be, but that’s not how it currently is. Looking at what’s available right now. You need 3 A100 40GB to get this amount of VRAM which will cost you way north of 20000$. Doing it with A6000s is still about 15k$. There’s not that many high VRAM options out there you know..
Re: Micron Kicks Off Production of HBM3E Memory
#37Earlier quoted context omitted.
Intel doesn't sell a lot of graphics cards whatsoever though. Be the first to offer 64GB of VRAM for under $1000 and that could change pretty fast.
Not without CUDA unfortunately.
Re: Micron Kicks Off Production of HBM3E Memory
#38Earlier quoted context omitted.
Not without CUDA unfortunately.
a lot of the true AI value is context window size limited, not compute limited.
What you are talking about is highly optimized inference using accelerators, batching and speculative decoding to achieve high throughout. Once you have that then compute is irrelevant except in terms of cost, but if all you have is a small consumer grade GPU you will be compute limited at the extreme limits of your context window.
Re: Micron Kicks Off Production of HBM3E Memory
#39Earlier quoted context omitted.
a lot of the true AI value is context window size limited, not compute limited.
Assuming 50 input tokens per second, you could still be waiting ten minutes for a full 32k token prompt. What you are talking about is highly optimized inference using accelerators, batching and speculative decoding to achieve high throughout. Once you have that then compute is irrelevant except in terms of cost, but if all you have is a small consumer grade GPU you will be compute limited at the extreme limits of yo…
I don't need long answers, I need by site specific knowledge base