Earlier quoted context omitted.
yes, I can't imagine the architecture of a supercomputer is the right one for LLM training. But maybe? If not, spending years to design and build a system for weather and nuke simulations and ending up doing something that's totally not made for these systems is kind of a mind bender. I can imagine the conversations that led to this: "we need a government owned LLM" "okay what do we need" "lots of compute power" "wel…
How are the optimal architectures for weather simulations and LLM training different?
I'm honestly not sure, and hoping somebody comments here and provides more information as I'm genuinely interested.
1 - https://openai.com/research/scaling-kubernetes-to-7500-nodes
OpenAI says their biggest jobs run on MPI so maybe a supercomputer would be better?