Live data from Hacker News

Nvidia Hopper GPU Architecture and H100 Accelerator

anandtech.com

171–180 of 183 posts

Re: Nvidia Hopper GPU Architecture and H100 Accelerator

#171
post #34
post #3

Sounds like we need some new training methods. If training could take place locally and asynchronously instead of globally through backpropagation, the amount of energy could probably be significantly reduced.

Trying to reduce energy consumption for ML like this is so silly.

That's like a person driving the Model T in 1908 saying "trying to reduce gas efficiency is so silly".

Why are people so dumb when it comes to planning for the future? Does it require a 1973 oil crisis to make people concerned about potential issues? Why can't people be preventative instead of reactive? Isn't the entire point of an engineer to optimize what they're building for the good of humanity?

Re: Nvidia Hopper GPU Architecture and H100 Accelerator

#172
post #99

Earlier quoted context omitted.

Oh I totally expect the size of models to grow along with whatever hardware can provide. I really do wonder how much more you could squeeze out of a full pod of gen2-H100's, obviously the model size would be ludicrous, but how far are we into the realm of dimishing returns. Your point about MoE architectures certainly sounds like the more _useful_ deployment, but the research seems to be pushing towards ludicrously l…

I agree! The models will definitely keep getting bigger, and MoEs are a part of that trend, sorry if that wasn’t clear. A pod of gen2-H100s might have 256 GPUs with 40 TB of total memory, and could easily run a 10T param model. So I think we are far from diminishing returns on the hardware side :) The model quality also continues to get better at scale. Re. reading material, I would take a look at DeepSpeed’s blog po…

[deleted]

Re: Nvidia Hopper GPU Architecture and H100 Accelerator

#173

NVidia and AMD datacenter GPUs continue to diverge between focusing on deep learning and traditional scientific computing respectively.

Scientific computer scientists always prefer Nvidia because of CUDA and much better development experience on Nvidia tools.

Re: Nvidia Hopper GPU Architecture and H100 Accelerator

#175
post #53

Earlier quoted context omitted.

50% sparsity and rated at 700W. The new DGX is 10kW!

I was recently researching how you'd host systems like this in a datacentre and was blown away to find out that you can cool 40kW in a single air cooled rack - this might be old news for many, but it was 2x or 3x what I expected! Glad I'm not paying the electricity bill :)

Supermicro will sell you 40 kW CPU TDP in a single air-cooled rack, not counting the rest of the servers.

Re: Nvidia Hopper GPU Architecture and H100 Accelerator

#176

The Tensor cores will be great for machine learning and the FP32/FP64 fantastic for HPC, but I'd be surprised if there were a lot of applications using both of these features at once. I wonder if there's room for a competitor to come in and sell another huge accelerator but with only one of these two features either at a lower price or with more performance? Perhaps the power density would be too high if everything w…

Graphcore's IPU is a machine learning variant on that. Power density seems to be ok. The CTO's talks (used to, I'm out of date) talk about dark silicon a lot.

I share your suspicion that fp64 and ML workloads are distinct but can see each running on the same cluster at different times.

Re: Nvidia Hopper GPU Architecture and H100 Accelerator

#177

NVidia and AMD datacenter GPUs continue to diverge between focusing on deep learning and traditional scientific computing respectively.

how so? what is this bad at for scientific computing?

FP16 performance is great, but the FP64 performance isn't terribly compelling. Scientific computing is generally FP64.

Re: Nvidia Hopper GPU Architecture and H100 Accelerator

#178
post #177

Earlier quoted context omitted.

how so? what is this bad at for scientific computing?

FP16 performance is great, but the FP64 performance isn't terribly compelling. Scientific computing is generally FP64.

If you're comparing it to the MI250 you're comparing it to 2 separate chips on a single card. This is fundamentally different and unless you have an ideal workload or have optimized appropriately, it's not going to hit anywhere near the peak FLOPS if you have data movement between chips.

Re: Nvidia Hopper GPU Architecture and H100 Accelerator

#179
post #155

Earlier quoted context omitted.

Yes.

And you think they shouldn't have that right because of social concerns like accumulation of wealth?

No, because they don’t own my identity! If my name is valuable I can will it to them. I can will my money to them. I just don’t think they should be able to endorse political candidates with my name after I’m dead, unless I specifically gave them that right by contract.

Re: Nvidia Hopper GPU Architecture and H100 Accelerator

#180

Off topic but I can't stand when corporations use actual people's names for their marketing who never gave them the permission to do so. For something like Shakespeare or Cicero I'm OK with it but Grace Hopper was alive in my lifetime, and even Tesla feels a little weird. What gives you the right to use that person's reputation to shill your product?

I generally agree with you, but in this case I suspect Grace Hopper would be honored by it and also impressed with the engineering here. It's not like they slapped her name on a soda can or something.

That’s not NVIDIA’s place to decide.
Post reply on HN