Live data from Hacker News

Nvidia Announces H100 NVL – Max Memory Server Card for Large Language Models

anandtech.com

1–10 of 111 posts

Re: Nvidia Announces H100 NVL – Max Memory Server Card for Large Language Models

#2
A bit underwhelming - H100 was announced at GTC 2022, and represented a huge stride over A100. But a year later, H100 is still not generally available at any public cloud I can find, and I haven't yet seen ML researchers reporting any use of H100.

The new "NVL" variant adds ~20% more memory per GPU by enabling the sixth HBM stack (previously only five out of six were used). Additionally, GPUs now come in pairs with 600GB/s bandwidth between the paired devices. However, the pair then uses PCIe as the sole interface to the rest of the system. This topology is an interesting hybrid of the previous DGX (put all GPUs onto a unified NVLink graph), and the more traditional PCIe accelerator cards (star topology of PCIe links, host CPU is the root node). Probably not an issue, I think PCIe 5.0 x16 is already fast enough to not bottleneck multi-GPU training too much.

Re: Nvidia Announces H100 NVL – Max Memory Server Card for Large Language Models

#3

A bit underwhelming - H100 was announced at GTC 2022, and represented a huge stride over A100. But a year later, H100 is still not generally available at any public cloud I can find, and I haven't yet seen ML researchers reporting any use of H100. The new "NVL" variant adds ~20% more memory per GPU by enabling the sixth HBM stack (previously only five out of six were used). Additionally, GPUs now come in pairs with 6…

It is interesting that hopper isn’t widely available yet.

I have seen some benchmarks from academia but nothing in the private sector.

I wonder if they thought they were moving too fast and wanted to milk amphere/ada as long as possible.

Not having any competition whatsoever means Nvidia can release what they like when they like.

Re: Nvidia Announces H100 NVL – Max Memory Server Card for Large Language Models

#4

A bit underwhelming - H100 was announced at GTC 2022, and represented a huge stride over A100. But a year later, H100 is still not generally available at any public cloud I can find, and I haven't yet seen ML researchers reporting any use of H100. The new "NVL" variant adds ~20% more memory per GPU by enabling the sixth HBM stack (previously only five out of six were used). Additionally, GPUs now come in pairs with 6…

It is interesting that hopper isn’t widely available yet. I have seen some benchmarks from academia but nothing in the private sector. I wonder if they thought they were moving too fast and wanted to milk amphere/ada as long as possible. Not having any competition whatsoever means Nvidia can release what they like when they like.

Why bother when you can get cryptobros paying way over MSRP for 3090s?

Re: Nvidia Announces H100 NVL – Max Memory Server Card for Large Language Models

#5
post #4

Earlier quoted context omitted.

It is interesting that hopper isn’t widely available yet. I have seen some benchmarks from academia but nothing in the private sector. I wonder if they thought they were moving too fast and wanted to milk amphere/ada as long as possible. Not having any competition whatsoever means Nvidia can release what they like when they like.

Why bother when you can get cryptobros paying way over MSRP for 3090s?

Not just cryptobros. A100s are the current top of the line and it’s hard to find them available on AWS and Lambda. Vast.AI has plenty if you trust renting from a stranger.

AMD really needs to pick up the pace and make a solid competitive offering in deep learning. They’re slowly getting there but they are at least 2 generations out.

Re: Nvidia Announces H100 NVL – Max Memory Server Card for Large Language Models

#6
post #4

Earlier quoted context omitted.

Why bother when you can get cryptobros paying way over MSRP for 3090s?

Not just cryptobros. A100s are the current top of the line and it’s hard to find them available on AWS and Lambda. Vast.AI has plenty if you trust renting from a stranger. AMD really needs to pick up the pace and make a solid competitive offering in deep learning. They’re slowly getting there but they are at least 2 generations out.

I would take a huge performance hit to just not deal with Nvidia drivers. Unless things have changed, it is still not really possible to operate on AMD hardware without a list of gotchas.

Re: Nvidia Announces H100 NVL – Max Memory Server Card for Large Language Models

#7
The really interesting upcoming LLM products are from AMD and Intel... with catches.

- The Intel Falcon Shores XPU is basically a big GPU that can use DDR5 DIMMS directly, hence it can fit absolutely enormous models into a single pool. But it has been delayed to 2025 :/

- AMD have not mentioned anything about the (not delayed) MI300 supporting DIMMs. If it doesn't, its capped to 128GB, and its being marketed as an HPC product like the MI200 anyway (which you basically cannot find on cloud services).

Nvidia also has some DDR5 grace CPUs, but the memory is embedded and I'm not sure how much of a GPU they have. Other startups (Tenstorrent, Cerebras, Graphcore and such) seemed to have underestimated the memory requirements of future models.

Re: Nvidia Announces H100 NVL – Max Memory Server Card for Large Language Models

#9
post #8

The TDP row in the comparison table must be in error. It shows the card with dual GH100 GPUs at 700W and the one with a single GH100 GPU at 700-800W ?!

That's the SXM version, used for instance in servers like the DGX. It's also faster than the PCIe variation.

Re: Nvidia Announces H100 NVL – Max Memory Server Card for Large Language Models

#10
post #6

Earlier quoted context omitted.

Not just cryptobros. A100s are the current top of the line and it’s hard to find them available on AWS and Lambda. Vast.AI has plenty if you trust renting from a stranger. AMD really needs to pick up the pace and make a solid competitive offering in deep learning. They’re slowly getting there but they are at least 2 generations out.

I would take a huge performance hit to just not deal with Nvidia drivers. Unless things have changed, it is still not really possible to operate on AMD hardware without a list of gotchas.

Its still basically impossible to find MI200s in the cloud.

On desktops, only the 7000 series is kinda competitive for AI in particular, and you have to go out of your way to get it running quick in PyTorch. The 6000 and 5000 series just weren't designed for AI.

Post reply on HN