Nvidia's v100 product page [0] says that it gets about 15 (single precision) - 125 ("deep learning") teraflop/s at 250-300 watts (joules per second). That means that if everything's as perfectly efficient as a marketing product page, it gets about 250/125-300/15 = 2-25 joules per teraflop, putting this model at about 0.6-8 terajoules.
A gallon of gasoline has about 120e6 joules [1] (though if you wanted to compare with burning it in a car, it's only 20-25% efficient at best [2] so it'd be fewer joules/gallon).
This model took the equivalent of about 5,000-67,000 gallons of gasoline at best and at ideal perfect energy efficiency. I get that openAI has made a decision not to be efficient with their dollars in order to see what's possible with future tech, but that means not being efficient with energy either, and it's getting kinda crazy. Sure, microsoft data centers aren't gasoline powered, so maybe it is closer to this ideal energy efficiency, and it's definitely going to be a better carbon footprint, but god damn it just seems wasteful.
Hell, the new A100 (again going off marketing materials [3], so at least it's apples to apples) could do it about 4x more efficiently. Is this research really worth what it costs, when waiting a year makes it that much more efficient?
[0] https://www.nvidia.com/en-us/data-center/v100/
[1] https://www.calculateme.com/energy/gallons-of-gas/to-joules/....
[2] https://en.wikipedia.org/wiki/Engine_efficiency#Gasoline_(pe...
[3] https://devblogs.nvidia.com/nvidia-ampere-architecture-in-de...