At the risk of doing original research, one thing I don’t see a lot of discussion on is that AI companies don’t train one model at a time. Typical engineers will have maybe 5-10 mid-size models training at once. Large automated hyperparameter grid searches might need ensembles of hundreds or thousands of training runs to compare loss curves etc... Most of these will turn out to be duds of course. Only one model gets released, and that one’s energy efficiency is (presumably) what’s reported.
So we might have to multiply the training numbers by the number of employees doing active research, times the number of models they like to keep in flight at any given time.