How do we know this is the largest ML facility in the field? Can anybody rule out the idea that I could go to GCP right now and launch 1500 nodes with 4 GPUs each?
The difference is probably the interconnect, no? What does GCP have between nodes?
According to the article, HPE/Cray 200gb ethernet. GCP A2 instances with 16 GPUs have 100gb ethernet. I'm not sure it's very important, would depend on the workload.