K8s with 1M nodes
bchess.github.io
K8s with 1M nodes
1–10 of 85 posts
Re: K8s with 1M nodes
#2[1] what is a node? Typically it is a synonym for "server". In some configurations HPC schedulers allow node sharing. Then we talk about order of 100k cores to be scheduled.
Re: K8s with 1M nodes
#3Typical large scale high performance computing clusters are at a size of 10k nodes (for instance Jupiter and SuperMUC in Germany) [1]. These centers are quite remarkably big buildings. I wonder how much 1M node single k8s clusters there are in the world right now. Most likely at the hyperscalers. [1] what is a node? Typically it is a synonym for "server". In some configurations HPC schedulers allow node sharing. Then…
Re: K8s with 1M nodes
#4This assumption is completely out of touch, and is especially funny when the goal is to build an extra large cluster.
Re: K8s with 1M nodes
#5“Perhaps my spiciest take from this entire project: most clusters don’t actually need the level of reliability and durability that etcd provides.” This assumption is completely out of touch, and is especially funny when the goal is to build an extra large cluster.
Re: K8s with 1M nodes
#6Typical large scale high performance computing clusters are at a size of 10k nodes (for instance Jupiter and SuperMUC in Germany) [1]. These centers are quite remarkably big buildings. I wonder how much 1M node single k8s clusters there are in the world right now. Most likely at the hyperscalers. [1] what is a node? Typically it is a synonym for "server". In some configurations HPC schedulers allow node sharing. Then…
I'm sure they mean actual servers / not just cores. Even in traditional HPC it isn't abstracted to the level of individual cores usually since most HPC jobs care about memory bandwidth - even with Infiniband or other techniques throughput / latency is much worse than on a single machine. Of course, multiple machines are connected (usually using MPI / Infiniband) but important to try to minimize communication between nodes where possible.
For AI workloads, they are running GPUs - so 10K+ cores on a single device so even less likely to be talking about cores here.
Re: K8s with 1M nodes
#7“Perhaps my spiciest take from this entire project: most clusters don’t actually need the level of reliability and durability that etcd provides.” This assumption is completely out of touch, and is especially funny when the goal is to build an extra large cluster.
etcd is also the entire point of k8s. that it's a single self-contained framework and doesn't require an external backer service. there is no kubernetes without etcd. much of the "secret sauce" of kubernetes is the "watch etcd" logic that "watches" desired state and does the cybernetic loop to bring the observed state adhere to the desired state.
https://github.com/k3s-io/kine is a reasonably adequate substitute for etcd. sqlite, MySQL, PostgreSQL can also be substituted in. Etcd is from the ground up built to be more scale-out reliable, and that rocks to have baked in. But given how easy it is to substitute etcd out, I feel like we are at least a little off if we're trying to say "etcd is also the entire point of k8s" (the APIserver is)
Re: K8s with 1M nodes
#8“Perhaps my spiciest take from this entire project: most clusters don’t actually need the level of reliability and durability that etcd provides.” This assumption is completely out of touch, and is especially funny when the goal is to build an extra large cluster.
etcd is also the entire point of k8s. that it's a single self-contained framework and doesn't require an external backer service. there is no kubernetes without etcd. much of the "secret sauce" of kubernetes is the "watch etcd" logic that "watches" desired state and does the cybernetic loop to bring the observed state adhere to the desired state.
Re: K8s with 1M nodes
#9Earlier quoted context omitted.
etcd is also the entire point of k8s. that it's a single self-contained framework and doesn't require an external backer service. there is no kubernetes without etcd. much of the "secret sauce" of kubernetes is the "watch etcd" logic that "watches" desired state and does the cybernetic loop to bring the observed state adhere to the desired state.
The API server is the thing. It so happens that the API server can mostly be a thin shell over etcd. But etcd itself while so common is not sacrosanct. https://github.com/k3s-io/kine is a reasonably adequate substitute for etcd. sqlite, MySQL, PostgreSQL can also be substituted in. Etcd is from the ground up built to be more scale-out reliable, and that rocks to have baked in. But given how easy it is to substitute e…
Re: K8s with 1M nodes
#10“Perhaps my spiciest take from this entire project: most clusters don’t actually need the level of reliability and durability that etcd provides.” This assumption is completely out of touch, and is especially funny when the goal is to build an extra large cluster.
etcd is also the entire point of k8s. that it's a single self-contained framework and doesn't require an external backer service. there is no kubernetes without etcd. much of the "secret sauce" of kubernetes is the "watch etcd" logic that "watches" desired state and does the cybernetic loop to bring the observed state adhere to the desired state.