Live data from Hacker News

Nvidia Announces H100 NVL – Max Memory Server Card for Large Language Models

anandtech.com

21–30 of 111 posts

Re: Nvidia Announces H100 NVL – Max Memory Server Card for Large Language Models

#21

A bit underwhelming - H100 was announced at GTC 2022, and represented a huge stride over A100. But a year later, H100 is still not generally available at any public cloud I can find, and I haven't yet seen ML researchers reporting any use of H100. The new "NVL" variant adds ~20% more memory per GPU by enabling the sixth HBM stack (previously only five out of six were used). Additionally, GPUs now come in pairs with 6…

It is interesting that hopper isn’t widely available yet. I have seen some benchmarks from academia but nothing in the private sector. I wonder if they thought they were moving too fast and wanted to milk amphere/ada as long as possible. Not having any competition whatsoever means Nvidia can release what they like when they like.

The question is, do they not have much production, or is OpenAI and Microsoft buying every single one they produce?

Re: Nvidia Announces H100 NVL – Max Memory Server Card for Large Language Models

#22
post #13

I wonder how soon we'll see something tailored specifically for local applications. Basically just tons of VRAM to be able to load large models, but not bleeding edge perf. And eGPU form factor, ideally.

I'm not a ML scientist my any means, but Perf seems as important as RAM from what I'm reading. Running prompts in internal chain of thought (eating up more TPU time) appears to give much better output.

Re: Nvidia Announces H100 NVL – Max Memory Server Card for Large Language Models

#24

I was wondering today if we would start to see the reverse of this. Small ASICS or some kind of optimized for LLM Gpu for desktop / or maybe even laptops of mobile. It is evident I think that LLM are here to stay and will be a major part of computing for a while. Getting this local, so we aren't reliant on clouds would be a huge boon for personal computing. Even if its a "worse" experience, being able to load up an L…

A couple of the big players are already looking at developing their own chips.

Re: Nvidia Announces H100 NVL – Max Memory Server Card for Large Language Models

#25
GPUs are going to be weird, underconfigured and overpriced until there is real competition.

Whether or not there is real competition depends entirely on whether Intels Arc line of GPUs stays in the market.

AMD strangely has decided not to compete. Its newest GPU the 7900 XTX is an extremely powerful card, close to the top of the line Nvidia RTX 4090 in raster performance.

If AMD had introduced it with an aggressively low price then then they could have wedged Nvidia, which is determinbed to exploit it's market dominance by squeezing the maximum money out of buyers.

Instead, AMD has decided to simply follow Nvidia in squeezing for maximum prices, with AM prices slightly behind Nvidia.

It's a strange decision from AMD who is well behind in market and apparently seems disinterested in increasing that market share by competing aggressively.

So a third player is needed - Intel - it's alot harder for three companies to sit on outrageously high prices for years rather than compete with each other for market share.

Re: Nvidia Announces H100 NVL – Max Memory Server Card for Large Language Models

#26

I was wondering today if we would start to see the reverse of this. Small ASICS or some kind of optimized for LLM Gpu for desktop / or maybe even laptops of mobile. It is evident I think that LLM are here to stay and will be a major part of computing for a while. Getting this local, so we aren't reliant on clouds would be a huge boon for personal computing. Even if its a "worse" experience, being able to load up an L…

Software/hardware co-evolution. Wouldn't be the first time we went down that road to good effect.

For anything that can be run remotely, it'll always be deployed and optimized server-side first. Higher utilization means more economy.

Then trickle down to local and end user devices if it makes sense.

Re: Nvidia Announces H100 NVL – Max Memory Server Card for Large Language Models

#27
post #14

Please give us consumer cards with more than 24GB VRAM, Nvidia. It was a slap in the face when the 4090 had the same memory capacity as the 3090. A6000 is 5000 dollars, ain't no hobbyist at home paying for that.

Nvidia don't want consumers using consumer GPUs for business.

If you are a business user then you must pay Nvidia gargantuan amounts of money.

This is the outcome of a market leader with no real competition - you pay much more for lower power than the consumer GPUs and you are forced into ujsing their business GPUs through software license restrictions on the drivers.

Re: Nvidia Announces H100 NVL – Max Memory Server Card for Large Language Models

#28

The really interesting upcoming LLM products are from AMD and Intel... with catches. - The Intel Falcon Shores XPU is basically a big GPU that can use DDR5 DIMMS directly, hence it can fit absolutely enormous models into a single pool. But it has been delayed to 2025 :/ - AMD have not mentioned anything about the (not delayed) MI300 supporting DIMMs. If it doesn't, its capped to 128GB, and its being marketed as an HP…

Grace can be paired with Hopper via a 900GB/s NVLINK bus (500GB/s memory bandwidth), 1TB of LPDDR5 on the CPU and 80-94GB of HBM3 on the GPU.

Re: Nvidia Announces H100 NVL – Max Memory Server Card for Large Language Models

#29
post #23

I'm super duper curious if there are ways to glob together VRAM between consumer-grade hardware to make this whole market more accessible to the common hacker?

You can, for instance, connect two RTX 3090 with an NVLink bridge. That gives you 48 GB in total. The 4090 doesn't support NVLink anymore.

Re: Nvidia Announces H100 NVL – Max Memory Server Card for Large Language Models

#30

I was wondering today if we would start to see the reverse of this. Small ASICS or some kind of optimized for LLM Gpu for desktop / or maybe even laptops of mobile. It is evident I think that LLM are here to stay and will be a major part of computing for a while. Getting this local, so we aren't reliant on clouds would be a huge boon for personal computing. Even if its a "worse" experience, being able to load up an L…

A couple of the big players are already looking at developing their own chips.

Have been for years. Maybe lots of years. It's expensive to have a go (many engineers plus cost of making the things) and it's difficult to beat the established players unless you see something they're doing wrong or your particular niche really cares about something the off the shelf hardware doesn't.
Post reply on HN