Live data from Hacker News

Nvidia Announces H100 NVL – Max Memory Server Card for Large Language Models

anandtech.com

101–110 of 111 posts

Re: Nvidia Announces H100 NVL – Max Memory Server Card for Large Language Models

#101
post #73
post #18

Earlier quoted context omitted.

The two cards show as two distinct GPUs to the host, connected via NVLink. Unification / load balancing happens via software.

Kinda depressing if you consider how they removed NVLink in the 4090, stating the following reason: > “The reason we took [NVLink] off is that we need I/O for other things, so we’re using that area to cram in as many AI processors as possible,” Jen-Hsun Huang explained of the reason for axing NVLink.[0] "NVLink is bad for your games and AI, trust me bro." But then this card, actually aimed at ML applications, uses it…

[deleted]

Re: Nvidia Announces H100 NVL – Max Memory Server Card for Large Language Models

#102
post #73
post #18

Earlier quoted context omitted.

The two cards show as two distinct GPUs to the host, connected via NVLink. Unification / load balancing happens via software.

Kinda depressing if you consider how they removed NVLink in the 4090, stating the following reason: > “The reason we took [NVLink] off is that we need I/O for other things, so we’re using that area to cram in as many AI processors as possible,” Jen-Hsun Huang explained of the reason for axing NVLink.[0] "NVLink is bad for your games and AI, trust me bro." But then this card, actually aimed at ML applications, uses it…

Market segmentation. Back when the Pascal architecture was the latest thing, it didn't make much sense to buy expensive Tesla P100 GPUs for many professional applications when consumer GeForce 1080 Ti cards gave you much more bang for the buck with few drawbacks. From the corporation's perspective it makes so much sense to differentiate the product lines more, now that their customers are deeply entrenched.

Re: Nvidia Announces H100 NVL – Max Memory Server Card for Large Language Models

#103

Earlier quoted context omitted.

Go with 2x 3090s instead. 4000 series doesn't support SLI, so you're stuck with the max of whatever one card you get.

If I remember correctly the NVLINK adds 100GB/s (where PCIE 4.0 is 64GB/s). Is it really worth getting 3090 performance (roughly half) for that extra bus speed?

Ampere NVLink (NV3) was 600 GByte/sec, with Hopper (NV4) it's 900 GByte/sec. https://www.nvidia.com/en-us/data-center/nvlink/

Re: Nvidia Announces H100 NVL – Max Memory Server Card for Large Language Models

#104
post #103

Earlier quoted context omitted.

If I remember correctly the NVLINK adds 100GB/s (where PCIE 4.0 is 64GB/s). Is it really worth getting 3090 performance (roughly half) for that extra bus speed?

Ampere NVLink (NV3) was 600 GByte/sec, with Hopper (NV4) it's 900 GByte/sec. https://www.nvidia.com/en-us/data-center/nvlink/

That is for the data center NVLINK, according to Wikipedia, for GA102 (3090) it is a 56.25GB/s bidirectional, yielding 112.5GB/s total bus bandwidth.

Re: Nvidia Announces H100 NVL – Max Memory Server Card for Large Language Models

#105
I was just saying to a colleague the day before this announcement that the inevitable consequence of the popularity of large language models will be GPUs with more memory.

Previously, GPUs were designed for gamers, and no game really "needs" more than 16 GB of VRAM. I've seen reviews of the A100 and H100 cards saying that the 80GB is ample for even the most demanding usage.

Now? Suddenly GPUs with 1 TB of memory could be immediately used, at scale, by deep-pocket customers happy to throw their entire wallets at NVIDIA.

This new H100 NVL model is a Frankenstein's monster stitched together from whatever they had lying around. It's a desperate move to corner the market early as possible. It's just the beginning, a preview of the times to come.

There will be a new digital moat, a new capitalist's empire, built upon on the scarcity of cards "big enough" to run models that nobody but a handful of megacorps can afford to train.

In fact, it won't be enough to restrict access by making the models expensive to train. The real moat will be models too expensive to run. Users will have to sign up, get API keys, and stand in line.

"Safe use of AI" my ass. Safe profits, more like. Safe monopolies, safe from competition.

Re: Nvidia Announces H100 NVL – Max Memory Server Card for Large Language Models

#106
post #103

Earlier quoted context omitted.

Ampere NVLink (NV3) was 600 GByte/sec, with Hopper (NV4) it's 900 GByte/sec. https://www.nvidia.com/en-us/data-center/nvlink/

That is for the data center NVLINK, according to Wikipedia, for GA102 (3090) it is a 56.25GB/s bidirectional, yielding 112.5GB/s total bus bandwidth.

Ah, that's true, thanks. It's the same type of NVLink as on the A40 GPU. https://images.nvidia.com/content/Solutions/data-center/a40/...

Re: Nvidia Announces H100 NVL – Max Memory Server Card for Large Language Models

#107

Earlier quoted context omitted.

We're NOT business users, we just want to run our own LLM at home. Given the size of LLMs, this should be possible with just a little bit of extra VRAM.

I don't think you understand though, they don't WANT you. They WANT the version of you who makes $150k+ a year and will splurge $5k on a Quadro. If they had trouble selling stock we would see this niche market get catered to.

That IS me. $5K is not enough to run an LLM at home (beyond the non-functional reduced quantization smaller models).

Re: Nvidia Announces H100 NVL – Max Memory Server Card for Large Language Models

#108

Earlier quoted context omitted.

I don't think you understand though, they don't WANT you. They WANT the version of you who makes $150k+ a year and will splurge $5k on a Quadro. If they had trouble selling stock we would see this niche market get catered to.

That IS me. $5K is not enough to run an LLM at home (beyond the non-functional reduced quantization smaller models).

Ahh yes, looks like I was too generous with my numbers, the new Quadro with 48GB VRAM is $7k, so you probably would need $14k and a Threadripper/Xeon/EPYC workstation because you won't have enough PCIE lanes/RAM/Memory Bandwidth otherwise.

So maybe more accurate is $200k+ a year and $20-30k on a workstation.

I grew up on $20k a year, the numbers in tech. are baffling!

Re: Nvidia Announces H100 NVL – Max Memory Server Card for Large Language Models

#109
post #89

Earlier quoted context omitted.

There will probably be Chinese options as well. China has an incentive to provide a domestic competitor due to deteriorating relations with the U.S.

They certainly will have to try, since nvidia is banned from exporting A100 and H100 chips.

They do ship A800 and H800 to China. H800 is the H100 with a much slower memory bandwidth. A800 is also a tiered down version of the A100

Re: Nvidia Announces H100 NVL – Max Memory Server Card for Large Language Models

#110
post #17

Does AMD have a chance here in the short term (say 24 months)?

AMD seems to be focusing on traditional HPC, they've got a ton of 64 bit flops in their recent commercial model. I expect their server GPUs are mostly for chasing supercomputer contracts, which can be pretty lucrative, while they cede model training to NVidia.

For now nvidia is a very dominant player for sure but in long run do you see it changing, with competition from Amd-xilinx, intel or potential AI hardware startups,why have the startups or other big players failed to make dent in nvidia's dominance ? considering how big this market will be in coming years there should have been significant investment made by other players but they seem to be incompetent in making even a competitive chip and nvidia which is already so ahead is running even more faster expanding its software ecosystem across various industries.
Post reply on HN