Interesting work! This is really an engineering achievement and I wish there was usable code. Real-time checkpointing seems like obviously the future to me, but it's going to be an easy-to-use, high-performance implementation that make that reality. One of the things I would like to have seen in the paper is a better analysis of simply checkpointing more often. It's briefly touched on: > It is infeasible to arbitrari…
Companies that can train models this big should hire people with HPC experience. They'd point out the need for storage clusters with high-speed interconnects. If they lack storage capabilities, I wonder why they're doing HPC like that. They clearly need the storage.
Example that BLOOM was trained on lists 100+GB of RAM per node and PB's of storage:
http://www.idris.fr/eng/jean-zay/cpu/jean-zay-cpu-hw-eng.htm...