Live data from Hacker News

Nvidia Announces H100 NVL – Max Memory Server Card for Large Language Models

anandtech.com

91–100 of 111 posts

Re: Nvidia Announces H100 NVL – Max Memory Server Card for Large Language Models

#91

Earlier quoted context omitted.

Absolutely not. Computers used to be extremely centralized and the decentralization revolution powered a ton of progress in both software development and hardware development. You can run many AI applications locally today that would have required a massive investment in hardware not all that long ago. It's just that the bleeding edge is still in that territory. One major optimization avenue is the improvement of the…

The bleeding edge will always be in that territory. It still requires a massive investment today to run AI applications locally to produce anywhere near as good results. People are spending upwards of $2000 for a GPU just to get decent results when it comes to image generation, many forgoing this entirely and just giving Google a monthly fee to use their hardware. Which is the point, decentralization will always be p…

Todays scraps are yesterdays state-of-the-art, and that's very logical and applies to far more than just AI applications. It's the way research and development result in products and the subsequent optimization. This has been true since the dawn of time in one form or another. At some point stone tools were high tech and next to affordably. Then it was bronze, then iron, and at some point we finally hit steam power. From there to the industrial revolution was a relatively short span and from there to electricity, electronics, solid state, computers, personal computers, mobile phones, smartphones and so on in ever decreasing steps.

If anything the steps are now so closely following each other that we have far more trouble tracking the societal changes and dealing with them than that we have a problem with the lag between technological advancement and its eventual commoditization.

Re: Nvidia Announces H100 NVL – Max Memory Server Card for Large Language Models

#92
post #81

Earlier quoted context omitted.

> For example on 24GB, Llama 30B runs only in 4bit mode and very slowly why do you think adding vram, but not cores will make it run faster?..

I've been told the 4 bit quantization slows it down, but don't quote me on this since I was unable to benchmark at 8 bit locally In any case, you're right it might not be as significant, however, the quality of the output increases with 8/16bit, and running 65B is completely impossible on 24GB

It's not impossible, there are several projects which load model layer by layer for execution from the disk or ram, but it will be much slower.

Re: Nvidia Announces H100 NVL – Max Memory Server Card for Large Language Models

#93
post #70
post #44

Earlier quoted context omitted.

Why is the sentiment here so much that LLMs will somehow be decentralized and run locally at some point? Has the story of the internet so far not been that centralization has pretty much always won?

Hackers want to run LLMs locally just because. It's not a mainstream thing.

It makes business sense as well. It doesn't make much sense to build an entire company around the idea that OpenAI's APIs are always available and you won't eventually get screwed. "Be careful of basing your business on top of another" and all that yadda yadda.

If you want to build a business around LLMs, it makes a lot of sense to be able to run the core service of what you want to offer on your own infrastructure instead of rely on a 3rd party that most likely doesn't give more than 1% care about you.

Re: Nvidia Announces H100 NVL – Max Memory Server Card for Large Language Models

#95
post #70

Earlier quoted context omitted.

Hackers want to run LLMs locally just because. It's not a mainstream thing.

It makes business sense as well. It doesn't make much sense to build an entire company around the idea that OpenAI's APIs are always available and you won't eventually get screwed. "Be careful of basing your business on top of another" and all that yadda yadda. If you want to build a business around LLMs, it makes a lot of sense to be able to run the core service of what you want to offer on your own infrastructure i…

Running LLMs on your own servers doesn't mean PCs which is what this thread is about. A100/H100 is fine for a business but people can't justify them for personal use.

Re: Nvidia Announces H100 NVL – Max Memory Server Card for Large Language Models

#97

Earlier quoted context omitted.

Not just cryptobros. A100s are the current top of the line and it’s hard to find them available on AWS and Lambda. Vast.AI has plenty if you trust renting from a stranger. AMD really needs to pick up the pace and make a solid competitive offering in deep learning. They’re slowly getting there but they are at least 2 generations out.

It's crazy to me that no other hardware company has sought to compete for the deep learning training/inference market yet ... The existing ecosystems (cuda, pytorch etc) are all pretty garbage anyway -- aside from the massive number of tutorials it doesn't seem like it would actually be hard to build a vertically integrated competitor ecosystem ... it feels a little like the rise of rails to me -- is a million articl…

No other company has sought this?

https://www.cerebras.net/ Has innovative technology, has actual customers, and is gaining a foothold in software-system stacks by integrating their platform into the OpenXLA GPU compiler.

Re: Nvidia Announces H100 NVL – Max Memory Server Card for Large Language Models

#99
post #89

Earlier quoted context omitted.

How could their moat possibly be deeper? First of all you need hardware with cutting-edge chips. Chips which can only be supplied by TSMC and Samsung. Then you need the software ranging all the way from the firmware and driver over something analogous to CUDA with libraries like cuDNN, cuBLAS and many others to integrations into pytorch and tensorflow. And none of that will come for free, like it came to Nvidia. Nvid…

There will probably be Chinese options as well. China has an incentive to provide a domestic competitor due to deteriorating relations with the U.S.

They certainly will have to try, since nvidia is banned from exporting A100 and H100 chips.

Re: Nvidia Announces H100 NVL – Max Memory Server Card for Large Language Models

#100

A bit underwhelming - H100 was announced at GTC 2022, and represented a huge stride over A100. But a year later, H100 is still not generally available at any public cloud I can find, and I haven't yet seen ML researchers reporting any use of H100. The new "NVL" variant adds ~20% more memory per GPU by enabling the sixth HBM stack (previously only five out of six were used). Additionally, GPUs now come in pairs with 6…

>H100 was announced at GTC 2022, and represented a huge stride over A100. But a year later, H100 is still not generally available at any public cloud I can find

You can safely assume an entity bought as many as they could.

Post reply on HN