Live data from Hacker News

Nvidia Hopper GPU Architecture and H100 Accelerator

anandtech.com

1–10 of 183 posts

Re: Nvidia Hopper GPU Architecture and H100 Accelerator

#6
post #3

Sounds like we need some new training methods. If training could take place locally and asynchronously instead of globally through backpropagation, the amount of energy could probably be significantly reduced.

The principled way of doing this is via ensemble learning, combining the predictions of multiple separately-trained models. But perhaps there are ways of improving that by including "global" training as well, where the "separate" models are allowed to interact while limiting overall training costs.
Post reply on HN