Disclosure: I used to work on Google Cloud. Perlmutter seems like an awesome system. But, I think the “ai exaflops” is a “X GPUS times the NVIDIA peak rate”. The new sparsity features on A100 are promising, but haven’t been demonstrated to be nearly as awesome in practice (yet). It also all comes down to workloads: large-scale distributed training is a funny workload! It’s not like LINPACK. If you make your model com…
Ultimately this is a machine to solve everybody’s needs, not just ML needs (although it must solve those needs very well in any case)
I’m not sure Cori will be remembered particularly fondly, but Perlmutter will probably be because it seems like it will be versatile enough to meet everyone’s needs.
For those wondering why not cloud - most simulations generally make sense in cloud because there’s nowhere near the data movement/storage involved and anything in particle physics is always embarrassingly parallel. The data isn’t moving there though - it is very cheap to keep the data at a place like NERSC especially with tape in the mix.