Live data from Hacker News

Nvidia DGX GH200 Whitepaper

resources.nvidia.com

1–10 of 45 posts

Re: Nvidia DGX GH200 Whitepaper

#2
Why is this called a whitepaper, as this is more of a documentation and architecture overview of the cluster? Wow a CLOS topology for networking, very innovative.

Details on NVLink would be great. For example, the needs and problems solved by their custom cables seemingly required by NVLink would be worth a whitepaper.

Don't get me wrong, this is still great the general public can get a glimpse into Grace Hopper. And they do a good job of simplifying while throwing around mind-boggling numbers (the NVLink bandwidth is insane, though no words on latency, crucial for remote memory access).

Re: Nvidia DGX GH200 Whitepaper

#7
post #4

So basically 2x faster than H100

They claim 1.1x to 7x, depending on what you're doing. The 10% to 50% is for the ~10k GPU LLM training, where the main bottleneck tends to be networking:

> DGX GH200 enables more efficient parallel mapping and alleviates the networking communication bottleneck. As a result, up to 1.5x faster training time can be achieved over a DGX H100-based solution for LLM training at scale.

Re: Nvidia DGX GH200 Whitepaper

#10
I would be interesting to know what kind of next-gen models this can train.

On the LLM frontier, we’re starting to hit the limits of reasoning abilities in the current gen.

Post reply on HN