Why is this called a whitepaper, as this is more of a documentation and architecture overview of the cluster? Wow a CLOS topology for networking, very innovative. Details on NVLink would be great. For example, the needs and problems solved by their custom cables seemingly required by NVLink would be worth a whitepaper. Don't get me wrong, this is still great the general public can get a glimpse into Grace Hopper. And…
Nvidia DGX GH200 Whitepaper
11–20 of 45 posts
Re: Nvidia DGX GH200 Whitepaper
#12I heard Elon say something interesting during the discussion/launch of xAI: "My prediction is that we will go from an extreme silicon shortage today, to probably a voltage-transformer shortage in about year, and then an electricity shortage in about a year, two years."
I'm not sure about the timeline, but it's an intriguing idea that soon the rate limiting resource will be electricity. I wonder how true that is and if we're prepared for that.
Re: Nvidia DGX GH200 Whitepaper
#13What's funny is that even though the DGX GH200 is some of the most powerful hardware available, there's such a voracious demand that it's not gonna be enough to quench it. In fact, this is one of those cases where I think the demand will always outpace supply. Exciting stuff ahead. I heard Elon say something interesting during the discussion/launch of xAI: "My prediction is that we will go from an extreme silicon sho…
To a first approximation, the amount of silicon wafers going through fabs globally is constant. We won’t suddenly increase chip manufacturing a hundredfold! There aren’t enough fabs or “tools” like the ASML EUV machines for that.
Electricity is used for lots of things, not just compute, and within compute the AI fraction is tiny. We’re ramping up a rounding error to a slightly larger rounding error.
What will increase is global energy demand for overall economic activity as manufacturing and industry is accelerated by AIs.
Anyone who’s played games like Factorio would know intuitively that the only two real inputs to the economy are raw materials and energy. Increases to manufacturing speed need matching increases to energy supply!
Re: Nvidia DGX GH200 Whitepaper
#14What's funny is that even though the DGX GH200 is some of the most powerful hardware available, there's such a voracious demand that it's not gonna be enough to quench it. In fact, this is one of those cases where I think the demand will always outpace supply. Exciting stuff ahead. I heard Elon say something interesting during the discussion/launch of xAI: "My prediction is that we will go from an extreme silicon sho…
He’s just plain wrong about the electricity usage going up because of AI compute. To a first approximation, the amount of silicon wafers going through fabs globally is constant. We won’t suddenly increase chip manufacturing a hundredfold! There aren’t enough fabs or “tools” like the ASML EUV machines for that. Electricity is used for lots of things, not just compute, and within compute the AI fraction is tiny. We’re…
Re: Nvidia DGX GH200 Whitepaper
#15What's funny is that even though the DGX GH200 is some of the most powerful hardware available, there's such a voracious demand that it's not gonna be enough to quench it. In fact, this is one of those cases where I think the demand will always outpace supply. Exciting stuff ahead. I heard Elon say something interesting during the discussion/launch of xAI: "My prediction is that we will go from an extreme silicon sho…
An Nvidia A100 costs $10000 and consumes 300W.
It seems unlikely that anyone could afford the number of A100s needed to create an electricity shortage.
If there is an electricity shortage, far more likely that ageing infrastructure and rising demand for air conditioning and electric car charging are to blame.
Re: Nvidia DGX GH200 Whitepaper
#16Re: Nvidia DGX GH200 Whitepaper
#17Why is this called a whitepaper, as this is more of a documentation and architecture overview of the cluster? Wow a CLOS topology for networking, very innovative. Details on NVLink would be great. For example, the needs and problems solved by their custom cables seemingly required by NVLink would be worth a whitepaper. Don't get me wrong, this is still great the general public can get a glimpse into Grace Hopper. And…
That’s what a marketing white paper is and does. It’s not an academic paper.
Re: Nvidia DGX GH200 Whitepaper
#18So basically 2x faster than H100
Re: Nvidia DGX GH200 Whitepaper
#19Earlier quoted context omitted.
He’s just plain wrong about the electricity usage going up because of AI compute. To a first approximation, the amount of silicon wafers going through fabs globally is constant. We won’t suddenly increase chip manufacturing a hundredfold! There aren’t enough fabs or “tools” like the ASML EUV machines for that. Electricity is used for lots of things, not just compute, and within compute the AI fraction is tiny. We’re…
A wafer of H100s uses far more electricity than a wafer of [Apple] A16s though.
However, that discounts the waste on the edges of the circular wafer, as well as the chip yield, which will both likely be worse for the larger chip [3]. But, assuming a generous 70% yield by area [4], one wafer's worth of H100s all packaged into GPUs and running full blast will use maybe 20 kilowatts, while the same wafer of A16s might use 3.6 kilowatts. Although in practice, the A16s will spend most of their time conserving battery power in your pocket, and even the H100s will spend some of their time idle.
TSMC is now producing over 14 million wafers per year. At most 1.2 million of those are on the 3nm node, and not all of that production goes to GPUs. But as an upper bound, if we imagine that all of TSMC's wafers could be filled up with nothing but H100 chips, and if all of those H100 chips were immediately put to use running AI 24/7, how much additional load could it put on the power grid every year?
The answer is, around 280 gigawatts, or if they were running 24/7 for a year, about 2500 terawatt-hours. That's about 10% of current world electricity consumption! So it's not completely implausible to imagine that a huge ramp-up in AI usage might have an effect on the electric grid.
*edit: This assumes we're talking about the Apple A16 (ie. the difference between phone chips and GPU chips). If we're talking about the Nvidia A16 (ie. the difference between current GPU chips and last node's GPU chips) see pclmulqdq's comment. ⠀
[1] https://nanoreview.net/en/soc/apple-a16-bionic
[2] https://www.techpowerup.com/gpu-specs/h100-pcie-80-gb.c3899
[3] https://news.ycombinator.com/item?id=24185108
[4] https://www.extremetech.com/computing/analyst-tsmc-hitting-5...
[5] https://www.tsmc.com/english/dedicatedFoundry/manufacturing/...
[6] https://www.wolframalpha.com/input?i=%2814+million%29+*+%282...*
Re: Nvidia DGX GH200 Whitepaper
#20Earlier quoted context omitted.
He’s just plain wrong about the electricity usage going up because of AI compute. To a first approximation, the amount of silicon wafers going through fabs globally is constant. We won’t suddenly increase chip manufacturing a hundredfold! There aren’t enough fabs or “tools” like the ASML EUV machines for that. Electricity is used for lots of things, not just compute, and within compute the AI fraction is tiny. We’re…
A wafer of H100s uses far more electricity than a wafer of [Apple] A16s though.
A16 is 200 sq mm of silicon while an H100 is about 800. That means you get about 100-120 A16's on a wafer, while you only get ~30 H100's (see https://www.silicon-edge.co.uk/j/index.php/resources/die-per...).
Let's assume yield is 100% to make things easier. The rated max power of the A16 is about 250W, while the H100 is quoted at 700W. Thus, a wafer of A16's is about 25-30 kW of power, while a wafer of H100's is about 21 kW.
Edit: Just clarifying, this is not about the Apple A16, but the Nvidia A16. The mobile process used by the Apple chips is built for much lower performance and power, so I can't imagine the two chips being anywhere near comparable - they fill two completely different roles.