Earlier quoted context omitted.
You can buy A100s in a server today, a number of integrators will happily sell it to you.
As someone who've tried for some weeks, it really seems like it's out-of-stock literally everywhere. The demand seems to be a lot higher than the supply at the moment, so much that I'm considering buying one myself instead of renting servers with it.
Nvidia Hopper GPU Architecture and H100 Accelerator
141–150 of 183 posts
Re: Nvidia Hopper GPU Architecture and H100 Accelerator
#142Earlier quoted context omitted.
It's a 32-bit format in memory and the additions are done with 32-bits.
I admit that I don't have the hardware to test your claims. But pretty much all the whitepapers I can find on TF32 explicitly state the 10-bit mantissa, suggesting that this is at best, a 19-bit format. 1-bit sign + 8-bit exponent + 10-bit mantissa. Yes, the system will read/write the 32-bit value to RAM. But if there's only 10-bits of mantissa in the circuits, you're only going to get 10-bits of precision (best case…
Re: Nvidia Hopper GPU Architecture and H100 Accelerator
#143Earlier quoted context omitted.
As someone who've tried for some weeks, it really seems like it's out-of-stock literally everywhere. The demand seems to be a lot higher than the supply at the moment, so much that I'm considering buying one myself instead of renting servers with it.
Does it make sense that all the GPUs are bought out? They each provide a return for mining in the short-term. In the long term, they can be used to run A(G)I models, which will be very very useful
Re: Nvidia Hopper GPU Architecture and H100 Accelerator
#144Sounds like we need some new training methods. If training could take place locally and asynchronously instead of globally through backpropagation, the amount of energy could probably be significantly reduced.
Disclosure: I work at MosaicML Yeah, I strongly agree. While Nvidia is working on better hardware (and they're doing a great job at it!), we believe that better training methods should be a big source of efficiency. We've released a new PyTorch library for efficient training at http://github.com/mosaicml/composer . Our combinations of methods can train CV models ~4x faster to the same accuracy on CV tasks, and ~2x fa…
Re: Nvidia Hopper GPU Architecture and H100 Accelerator
#145Earlier quoted context omitted.
Does it make sense that all the GPUs are bought out? They each provide a return for mining in the short-term. In the long term, they can be used to run A(G)I models, which will be very very useful
This is the GPU the parent is talking about https://www.nvidia.com/en-us/data-center/a100/
Re: Nvidia Hopper GPU Architecture and H100 Accelerator
#146Re: Nvidia Hopper GPU Architecture and H100 Accelerator
#147Earlier quoted context omitted.
Did you check Lambda or Exxact?
Yes, nor Lambda Labs or Exxact Corporation have them available last time I checked (last week). Both citing high demand as the reason for it being unavailable.
If you're interested in giving us a shot, feel free to shoot me an email at mike at crusoecloud dot com.
Re: Nvidia Hopper GPU Architecture and H100 Accelerator
#148Earlier quoted context omitted.
The main tensor op is a matmul intrinsic which is useful for way more than just deep learning. Edit; many of these speeds are low precision which is less useful outside of deep learning, but the higher precision matmul ops in the tensor cores are still very fast and very useful for wide variety of tasks.
> but the higher precision matmul ops in the tensor cores are still very fast and very useful for wide variety of tasks. The FP64 matrix-multiplication is only 60 TFlops, no where near the advertized 1000 TFlops. TF32 matrix-multiplication is a poorly named 16-bit operation.
Re: Nvidia Hopper GPU Architecture and H100 Accelerator
#1491 petaflop on a chip?? What is the catch?
700 watts so being NVidia it'll blow up in 6 months and you'll need to wait in a queue for 6 months to RMA it because all the miners had bought up the entire supply chain.