Live data from Hacker News

Nvidia Hopper GPU Architecture and H100 Accelerator

anandtech.com

21–30 of 183 posts

Re: Nvidia Hopper GPU Architecture and H100 Accelerator

#21

Earlier quoted context omitted.

Given that it's Nvidia, no Linux support. That's the catch.

I thought that only applied to their consumer products.

Their consumer products have Linux support too, the catch is just that the drivers are proprietary binary blobs

Re: Nvidia Hopper GPU Architecture and H100 Accelerator

#22

Earlier quoted context omitted.

Given that it's Nvidia, no Linux support. That's the catch.

I thought that only applied to their consumer products.

Don't they provide Linux drivers for their gaming graphics cards too, just not open source?

Re: Nvidia Hopper GPU Architecture and H100 Accelerator

#23

This seems fast... TF32 ....... 1,000 TFLOPS (tensor core) FP64/FP32 ... 60 TFLOPS I am more interested in the 144-core Grace CPU Superchip. nVidia is getting into the CPU business...

I think the 1PFLOPS figure for TF32 is with sparsity, which should be called out in the name. Maybe ‘TFS32’? I mainly use dense FP16 so the 1PFLOPS for that looks pretty good.

Re: Nvidia Hopper GPU Architecture and H100 Accelerator

#25

Earlier quoted context omitted.

Given that it's Nvidia, no Linux support. That's the catch.

Nvidia provides Linux drivers for their server chips.

Don't they provide them for their consumer cards too, just that it's a closed source binary blob?

Re: Nvidia Hopper GPU Architecture and H100 Accelerator

#26
"Combined with the additional memory on H100 and the faster NVLink 4 I/O, and NVIDIA claims that a large cluster of GPUs can train a transformer up to 9x faster, which would bring down training times on today’s largest models down to a more reasonable period of time, and make even larger models more practical to tackle."

Looking good.

Re: Nvidia Hopper GPU Architecture and H100 Accelerator

#29

"Combined with the additional memory on H100 and the faster NVLink 4 I/O, and NVIDIA claims that a large cluster of GPUs can train a transformer up to 9x faster, which would bring down training times on today’s largest models down to a more reasonable period of time, and make even larger models more practical to tackle." Looking good.

We have needed wide use of NVlink or something like it for a long time now......heres to hoping mobo manufacturers actually widely implement it!

Re: Nvidia Hopper GPU Architecture and H100 Accelerator

#30

1 petaflop on a chip?? What is the catch?

Tensor petaflops are useful in only very few circumstances. One of which is the highly lucrative deep learning community though.

The main tensor op is a matmul intrinsic which is useful for way more than just deep learning.

Edit; many of these speeds are low precision which is less useful outside of deep learning, but the higher precision matmul ops in the tensor cores are still very fast and very useful for wide variety of tasks.

Post reply on HN