I was wondering today if we would start to see the reverse of this. Small ASICS or some kind of optimized for LLM Gpu for desktop / or maybe even laptops of mobile. It is evident I think that LLM are here to stay and will be a major part of computing for a while. Getting this local, so we aren't reliant on clouds would be a huge boon for personal computing. Even if its a "worse" experience, being able to load up an L…
Why is the sentiment here so much that LLMs will somehow be decentralized and run locally at some point? Has the story of the internet so far not been that centralization has pretty much always won?
Nvidia Announces H100 NVL – Max Memory Server Card for Large Language Models
51–60 of 111 posts
Re: Nvidia Announces H100 NVL – Max Memory Server Card for Large Language Models
#52I would sell a kidney for one of these. It's basically impossible to train language models on a consumer 24GB card. The jump up is the A6000 ADA, at 48GB for $8,000. This one will probably be priced somewhere in the $100k+ range.
You think? It’s double 48 GB (per card) so why wouldn’t it be in the $20k range?
Re: Nvidia Announces H100 NVL – Max Memory Server Card for Large Language Models
#53I was wondering today if we would start to see the reverse of this. Small ASICS or some kind of optimized for LLM Gpu for desktop / or maybe even laptops of mobile. It is evident I think that LLM are here to stay and will be a major part of computing for a while. Getting this local, so we aren't reliant on clouds would be a huge boon for personal computing. Even if its a "worse" experience, being able to load up an L…
Why is the sentiment here so much that LLMs will somehow be decentralized and run locally at some point? Has the story of the internet so far not been that centralization has pretty much always won?
Re: Nvidia Announces H100 NVL – Max Memory Server Card for Large Language Models
#54Earlier quoted context omitted.
>> The root cause is that TSMC raised prices in everyone. This is not correct.
https://www.tomshardware.com/news/tsmc-ups-chip-production-p...
You are incorrect that this is the root cause of GPU prices being sky high.
If manufacturing cost was the root cause then it would be simply impossible to bring prices down without losing money.
The root cause of GPU prices being so high is lack of competition - AMD and Nvidia are choosing to maximise profit, and they are deliberately undersupplying the market to create scarcity and therefore prop up prices.
"AMD 'undershipping' chips to help prop prices up" https://www.pcgamer.com/amd-undershipping-chips-to-help-prop...
"AMD is ‘undershipping’ chips to balance CPU, GPU supply Less supply to balance out demand—and keep prices high." https://www.pcworld.com/article/1499957/amd-is-undershipping...
In summary, GPOU prices are ridiculously high because Nvidia and AMD are overpricing them because they believe this is what gamers will pay, NOT because manufacturing costs have forced prices to be high.
Re: Nvidia Announces H100 NVL – Max Memory Server Card for Large Language Models
#55I would sell a kidney for one of these. It's basically impossible to train language models on a consumer 24GB card. The jump up is the A6000 ADA, at 48GB for $8,000. This one will probably be priced somewhere in the $100k+ range.
Use 4 consumer grade 4090 then. It would be much cheaper and better in almost every aspect. Also even with this, forget about training foundational models. Meta spent 82k GPU hours on the smallest llama and 1M hours on largest.
Re: Nvidia Announces H100 NVL – Max Memory Server Card for Large Language Models
#56The really interesting upcoming LLM products are from AMD and Intel... with catches. - The Intel Falcon Shores XPU is basically a big GPU that can use DDR5 DIMMS directly, hence it can fit absolutely enormous models into a single pool. But it has been delayed to 2025 :/ - AMD have not mentioned anything about the (not delayed) MI300 supporting DIMMs. If it doesn't, its capped to 128GB, and its being marketed as an HP…
Grace can be paired with Hopper via a 900GB/s NVLINK bus (500GB/s memory bandwidth), 1TB of LPDDR5 on the CPU and 80-94GB of HBM3 on the GPU.
Re: Nvidia Announces H100 NVL – Max Memory Server Card for Large Language Models
#57The really interesting upcoming LLM products are from AMD and Intel... with catches. - The Intel Falcon Shores XPU is basically a big GPU that can use DDR5 DIMMS directly, hence it can fit absolutely enormous models into a single pool. But it has been delayed to 2025 :/ - AMD have not mentioned anything about the (not delayed) MI300 supporting DIMMs. If it doesn't, its capped to 128GB, and its being marketed as an HP…
> DDR5 DIMMS directly That's the problem. Good DDR5 RAM's memory speed is <100GB/s, while nvidia could has up to 2TB/s, and still the bottleneck lies on memory speed for most applications.
Anyway, what I was implying is that simply fitting a trillion parameter model into a single pool is probably more efficient than splitting it up over a power hungry interconnect. Bandwidth is much lower, but latency is also slower, you are shuffling much less data around.
Re: Nvidia Announces H100 NVL – Max Memory Server Card for Large Language Models
#58Please give us consumer cards with more than 24GB VRAM, Nvidia. It was a slap in the face when the 4090 had the same memory capacity as the 3090. A6000 is 5000 dollars, ain't no hobbyist at home paying for that.
Nvidia don't want consumers using consumer GPUs for business. If you are a business user then you must pay Nvidia gargantuan amounts of money. This is the outcome of a market leader with no real competition - you pay much more for lower power than the consumer GPUs and you are forced into ujsing their business GPUs through software license restrictions on the drivers.
Re: Nvidia Announces H100 NVL – Max Memory Server Card for Large Language Models
#59Please give us consumer cards with more than 24GB VRAM, Nvidia. It was a slap in the face when the 4090 had the same memory capacity as the 3090. A6000 is 5000 dollars, ain't no hobbyist at home paying for that.
Nvidia don't want consumers using consumer GPUs for business. If you are a business user then you must pay Nvidia gargantuan amounts of money. This is the outcome of a market leader with no real competition - you pay much more for lower power than the consumer GPUs and you are forced into ujsing their business GPUs through software license restrictions on the drivers.
Given the size of LLMs, this should be possible with just a little bit of extra VRAM.
Re: Nvidia Announces H100 NVL – Max Memory Server Card for Large Language Models
#60Earlier quoted context omitted.
Nvidia don't want consumers using consumer GPUs for business. If you are a business user then you must pay Nvidia gargantuan amounts of money. This is the outcome of a market leader with no real competition - you pay much more for lower power than the consumer GPUs and you are forced into ujsing their business GPUs through software license restrictions on the drivers.
That was always why the Titan line was so great - they typically unlocked features in between Quadro and Gaming cards. Sometimes it was subtle (like very good FP32 AND FP16 performance) or adding full 10 bit colour support if you had a Titan only. Now it seems like they have opened up even more of those features to consumer cards (at least the creative ones) with the studio drivers.