Why is this called a whitepaper, as this is more of a documentation and architecture overview of the cluster? Wow a CLOS topology for networking, very innovative. Details on NVLink would be great. For example, the needs and problems solved by their custom cables seemingly required by NVLink would be worth a whitepaper. Don't get me wrong, this is still great the general public can get a glimpse into Grace Hopper. And…
Nvidia DGX GH200 Whitepaper
31–40 of 45 posts
Re: Nvidia DGX GH200 Whitepaper
#32What's funny is that even though the DGX GH200 is some of the most powerful hardware available, there's such a voracious demand that it's not gonna be enough to quench it. In fact, this is one of those cases where I think the demand will always outpace supply. Exciting stuff ahead. I heard Elon say something interesting during the discussion/launch of xAI: "My prediction is that we will go from an extreme silicon sho…
Re: Nvidia DGX GH200 Whitepaper
#33The memory and bandwidth numbers are mind blowing. Going to be very hard to catch Nvidia. It’s as if competitors are going through the motions for participation prizes.
AMD has been shipping 128x lanes of PCIe 5.0 on chip. That's 0.5TBps. Getting up to 0.9TBps isn't that crazy, but having big enough fabric & switches to attach to is a huge feat. I have hope though. CXL switching is going to give the whole industry a very fresh look at interconnect fabrics, as a simpler to manage faster more direct alternative to PCIe. Should be good. Personally I worry it's flogging a dead horse, ha…
Re: Nvidia DGX GH200 Whitepaper
#34What's funny is that even though the DGX GH200 is some of the most powerful hardware available, there's such a voracious demand that it's not gonna be enough to quench it. In fact, this is one of those cases where I think the demand will always outpace supply. Exciting stuff ahead. I heard Elon say something interesting during the discussion/launch of xAI: "My prediction is that we will go from an extreme silicon sho…
Are there any examples at all about that guy being right about a technology prediction?
Re: Nvidia DGX GH200 Whitepaper
#35Earlier quoted context omitted.
Adding up to "1 exaFLOPS" (sparse FP8). For reference, the fastest FP64 supercomputer is the AMD-based Frontier supercomputer, at 1.1 exaFLOPS.
Does sparse mean anything other than we can not actually do as many FP8 operations per second as we just claimed? To me it sounds like they can do X matrix operations per second on sparse matrices using Y FP8 operations per second, but instead of just saying what Y is they tell us how many FP8 operations would be required if the matrices were not sparse. Is this pure marketing bullshit or is there some logic to this?…
Re: Nvidia DGX GH200 Whitepaper
#36Earlier quoted context omitted.
Are there any examples at all about that guy being right about a technology prediction?
Electric cars, rocket ships …
Given we're talking about hardware for software, let's at least look at his track record in the software industry... glances at Twitter ah, yeah, not great either.
Voltage regulator and electricity shortage from AI growth straight up doesn't make sense, it's dumber than the stuff he was spouting while "deep-diving" his Twitter misacquisition.
Re: Nvidia DGX GH200 Whitepaper
#37As context: 1x dgx gh200 has 256x gh200s which each have 1x h100 and 1x grace cpu
Adding up to "1 exaFLOPS" (sparse FP8). For reference, the fastest FP64 supercomputer is the AMD-based Frontier supercomputer, at 1.1 exaFLOPS.
Re: Nvidia DGX GH200 Whitepaper
#38Earlier quoted context omitted.
A wafer of H100s uses far more electricity than a wafer of [Apple] A16s though.
An H100 uses up to 350 Watts, while an A16 has a TDP of only 8 W. But, the A16 is a smaller chip (about 108mm vs. the H100's 814mm) so you can fit more of them on a wafer. Since a wafer is 300mm in diameter, its area is 70685 mm^2, which would yield 86 H100's or 654 A16's. [1][2] However, that discounts the waste on the edges of the circular wafer, as well as the chip yield, which will both likely be worse for the la…
1.2 x 30 x 30000($/board) ~ 1 trillion $$$. Time for NVDA call.
Re: Nvidia DGX GH200 Whitepaper
#39Earlier quoted context omitted.
An H100 uses up to 350 Watts, while an A16 has a TDP of only 8 W. But, the A16 is a smaller chip (about 108mm vs. the H100's 814mm) so you can fit more of them on a wafer. Since a wafer is 300mm in diameter, its area is 70685 mm^2, which would yield 86 H100's or 654 A16's. [1][2] However, that discounts the waste on the edges of the circular wafer, as well as the chip yield, which will both likely be worse for the la…
8 watts for the A16's TDP cannot be correct. Your phone CPU has a higher TDP. I saw 250 on Nvidia's website as a maximum. Edit: Oh, you are talking about the Apple A16. Those chips are completely different in function, so sure.
A few mobile phone chips had a higher TDP, up to 10 W, but those were notorious for overheating and for low battery life.
Re: Nvidia DGX GH200 Whitepaper
#40What's funny is that even though the DGX GH200 is some of the most powerful hardware available, there's such a voracious demand that it's not gonna be enough to quench it. In fact, this is one of those cases where I think the demand will always outpace supply. Exciting stuff ahead. I heard Elon say something interesting during the discussion/launch of xAI: "My prediction is that we will go from an extreme silicon sho…
He’s just plain wrong about the electricity usage going up because of AI compute. To a first approximation, the amount of silicon wafers going through fabs globally is constant. We won’t suddenly increase chip manufacturing a hundredfold! There aren’t enough fabs or “tools” like the ASML EUV machines for that. Electricity is used for lots of things, not just compute, and within compute the AI fraction is tiny. We’re…