Live data from Hacker News

Nvidia Announces H100 NVL – Max Memory Server Card for Large Language Models

anandtech.com

61–70 of 111 posts

Re: Nvidia Announces H100 NVL – Max Memory Server Card for Large Language Models

#61
post #34

What exactly is an SXM5 socket? It sounds like a PCIe competitor, but proprietary to nvidia. Looking at it, it seems specific to nvidia DGX (mother?)boards. Is this just a "better" alternative to PCIe (with power delivery, and such), or fundamentally a new technology?

Yes to all your questions. It's specifically designed for commercial compute servers. It provides significantly more bandwidth and speed over PCIe.

It's also enormously more expensive and I'm not sure if you can buy it new without getting the nvidia compute server.

Re: Nvidia Announces H100 NVL – Max Memory Server Card for Large Language Models

#62
post #29
post #23

I'm super duper curious if there are ways to glob together VRAM between consumer-grade hardware to make this whole market more accessible to the common hacker?

You can, for instance, connect two RTX 3090 with an NVLink bridge. That gives you 48 GB in total. The 4090 doesn't support NVLink anymore.

> The 4090 doesn't support NVLink anymore.

Are you sure about that?

Re: Nvidia Announces H100 NVL – Max Memory Server Card for Large Language Models

#63

Earlier quoted context omitted.

Nvidia don't want consumers using consumer GPUs for business. If you are a business user then you must pay Nvidia gargantuan amounts of money. This is the outcome of a market leader with no real competition - you pay much more for lower power than the consumer GPUs and you are forced into ujsing their business GPUs through software license restrictions on the drivers.

That was always why the Titan line was so great - they typically unlocked features in between Quadro and Gaming cards. Sometimes it was subtle (like very good FP32 AND FP16 performance) or adding full 10 bit colour support if you had a Titan only. Now it seems like they have opened up even more of those features to consumer cards (at least the creative ones) with the studio drivers.

Hmmm ... "Studio Drivers" ... how are these tangibly different to gaming drivers?

According to this, the difference seems to be that Studio Drivers are older and better tested, nothing else.

https://nvidia.custhelp.com/app/answers/detail/a_id/4931/~/n...

What am I missing in my understanding of Studio Drivers?

""" How do Studio Drivers differ from Game Ready Drivers (GRD)?

In 2014, NVIDIA created the Game Ready Driver program to provide the best day-0 gaming experience. In order to accomplish this, the release cadence for Game Ready Drivers is driven by the release of major new game content giving our driver team as much time as possible to work on a given title. In similar fashion, NVIDIA now offers the Studio Driver program. Designed to provide the ultimate in functionality and stability for creative applications, Studio Drivers provide extensive testing against top creative applications and workflows for the best performance possible, and support any major creative app updates to ensure that you are ready to update any apps on Day 1. ""

Re: Nvidia Announces H100 NVL – Max Memory Server Card for Large Language Models

#64

A bit underwhelming - H100 was announced at GTC 2022, and represented a huge stride over A100. But a year later, H100 is still not generally available at any public cloud I can find, and I haven't yet seen ML researchers reporting any use of H100. The new "NVL" variant adds ~20% more memory per GPU by enabling the sixth HBM stack (previously only five out of six were used). Additionally, GPUs now come in pairs with 6…

Yes, I was expecting a RAM-doubled edition of the H100, this is just a higher-binned version of the same part. I got an email from vultr, saying that they're "officially taking reservations for the NVIDIA HGX H100", so I guess all public clouds are going to get those soon.

[deleted]

Re: Nvidia Announces H100 NVL – Max Memory Server Card for Large Language Models

#65

Earlier quoted context omitted.

Nvidia don't want consumers using consumer GPUs for business. If you are a business user then you must pay Nvidia gargantuan amounts of money. This is the outcome of a market leader with no real competition - you pay much more for lower power than the consumer GPUs and you are forced into ujsing their business GPUs through software license restrictions on the drivers.

We're NOT business users, we just want to run our own LLM at home. Given the size of LLMs, this should be possible with just a little bit of extra VRAM.

Exactly, we're just below that sweet spot right now.

For example on 24GB, Llama 30B runs only in 4bit mode and very slowly, but I can imagine a RLHF finetuned 30B or 65B version running in at least 8bit would be actually useful, and you could run it on your own computer easily.

Re: Nvidia Announces H100 NVL – Max Memory Server Card for Large Language Models

#66
post #14

Please give us consumer cards with more than 24GB VRAM, Nvidia. It was a slap in the face when the 4090 had the same memory capacity as the 3090. A6000 is 5000 dollars, ain't no hobbyist at home paying for that.

Nvidia can't do a large 'consumer' card without cannibalizing their commercial ML business. ATI doesn't have that problem.

ATI seems to be holding the idiot ball.

Port stable diffusion and clip to their hardware. Train an upsized version sized for a 48GB card. Release a prosumer 48gb card... get huge uptake from artists and creators using the tech.

Re: Nvidia Announces H100 NVL – Max Memory Server Card for Large Language Models

#67
post #29

Earlier quoted context omitted.

You can, for instance, connect two RTX 3090 with an NVLink bridge. That gives you 48 GB in total. The 4090 doesn't support NVLink anymore.

> The 4090 doesn't support NVLink anymore. Are you sure about that?

that's what the press said: https://www.tomshardware.com/news/gigabyte-leaves-nvlink-tra...

Re: Nvidia Announces H100 NVL – Max Memory Server Card for Large Language Models

#68

Earlier quoted context omitted.

Not just cryptobros. A100s are the current top of the line and it’s hard to find them available on AWS and Lambda. Vast.AI has plenty if you trust renting from a stranger. AMD really needs to pick up the pace and make a solid competitive offering in deep learning. They’re slowly getting there but they are at least 2 generations out.

It's crazy to me that no other hardware company has sought to compete for the deep learning training/inference market yet ... The existing ecosystems (cuda, pytorch etc) are all pretty garbage anyway -- aside from the massive number of tutorials it doesn't seem like it would actually be hard to build a vertically integrated competitor ecosystem ... it feels a little like the rise of rails to me -- is a million articl…

There are tons of companies trying; they just aren't succeeding.

Re: Nvidia Announces H100 NVL – Max Memory Server Card for Large Language Models

#69
post #29
post #23

I'm super duper curious if there are ways to glob together VRAM between consumer-grade hardware to make this whole market more accessible to the common hacker?

You can, for instance, connect two RTX 3090 with an NVLink bridge. That gives you 48 GB in total. The 4090 doesn't support NVLink anymore.

You actually can split a model [0] onto multiple GPUs even without NVLink, just using the PCIe for the transfers.

Depending on the model the performance is sometimes not all that different. I believe for solely inference on some models the speed difference may barely be noticeable, where for other training activities it may make 10+% difference [1]

[0] https://pytorch.org/tutorials/intermediate/model_parallel_tu...

[1] https://huggingface.co/transformers/v4.9.2/performance.html

Re: Nvidia Announces H100 NVL – Max Memory Server Card for Large Language Models

#70
post #44

I was wondering today if we would start to see the reverse of this. Small ASICS or some kind of optimized for LLM Gpu for desktop / or maybe even laptops of mobile. It is evident I think that LLM are here to stay and will be a major part of computing for a while. Getting this local, so we aren't reliant on clouds would be a huge boon for personal computing. Even if its a "worse" experience, being able to load up an L…

Why is the sentiment here so much that LLMs will somehow be decentralized and run locally at some point? Has the story of the internet so far not been that centralization has pretty much always won?

Hackers want to run LLMs locally just because. It's not a mainstream thing.
Post reply on HN