It seems all cutting edge datacenters like x.ai Colossus are using Nvidia networking. Now Google is upgrading to Nvidia networking, too. Since Nvidia owns most of the Gpgpu products, they have top notch networking and interconnect, I wonder if they don't have a plan to own all datacenter hardware in the future. Maybe they plan to also release CPUs, motherboards, storage and whatever else is needed.
Yeah there’s a bit of industry worry about that very eventuality — hence the ultra Ethernet consortium trying to work on open source alternatives to the mellanox/nvidia lock-in. https://ultraethernet.org/
Speed, scale and reliability: 25 years of Google datacenter networking evolution
61–70 of 91 posts
Re: Speed, scale and reliability: 25 years of Google datacenter networking evolution
#62Earlier quoted context omitted.
What if you have dozens of big data centers?
To reinforce your point: The scale of cloud data centres reflects the scale of their customer base, not the size of the basket for each individual customer. Larger data centres actually improve availability through several mechanisms: more power components such as generators means the failure of any one is just a few percent instead of a total blackout. You can also partition core infrastructure like routers and powe…
Re: Speed, scale and reliability: 25 years of Google datacenter networking evolution
#63Earlier quoted context omitted.
To reinforce your point: The scale of cloud data centres reflects the scale of their customer base, not the size of the basket for each individual customer. Larger data centres actually improve availability through several mechanisms: more power components such as generators means the failure of any one is just a few percent instead of a total blackout. You can also partition core infrastructure like routers and powe…
I provided three different references. Despite the massive downvotes on my comment I guess by Google engineers, as a troll...:-)I take comfort on the fact nobody was able to advance a reference to prove me wrong.
It is true that the nomenclature "AWS Availability Zone" has a different meaning than "GCP Zone" when discussing the physical separation between zones within the same region.
It's unclear why this is inherently a bad thing, as long as them same overall level of reliability is achieved.
Re: Speed, scale and reliability: 25 years of Google datacenter networking evolution
#64Earlier quoted context omitted.
I provided three different references. Despite the massive downvotes on my comment I guess by Google engineers, as a troll...:-)I take comfort on the fact nobody was able to advance a reference to prove me wrong.
You haven't actually made an argument. It is true that the nomenclature "AWS Availability Zone" has a different meaning than "GCP Zone" when discussing the physical separation between zones within the same region. It's unclear why this is inherently a bad thing, as long as them same overall level of reliability is achieved.
Re: Speed, scale and reliability: 25 years of Google datacenter networking evolution
#65Pretty crazy. Supporting 1.5mbps video calls for each human on earth? Did I read that right? Just goes to show how drastic and extraordinary levels of scale can be.
Re: Speed, scale and reliability: 25 years of Google datacenter networking evolution
#66Earlier quoted context omitted.
You haven't actually made an argument. It is true that the nomenclature "AWS Availability Zone" has a different meaning than "GCP Zone" when discussing the physical separation between zones within the same region. It's unclear why this is inherently a bad thing, as long as them same overall level of reliability is achieved.
The phrase "as long as the same overall level of reliability is achieved" is logically flawed when discussing physically co-located vs. geographically separated infrastructure.
In my experience, the set of issues that would affect 2 buildings close to each other, but not two buildings a mile apart, is vanishingly small, usually just last mile fiber cuts or power issues (which are rare and mitigated by having multiple independent providers), as well as issues like building fires (which are exceedingly rare, we know of, perhaps two of notable impact in more than a decade across the big three cloud providers).
Everything else is done at the zone level no matter what (onsite repair work, rollouts, upgrades, control plane changes, etc.) or can impact an entire region (non-last mile fiber or power cuts, inclement weather, regional power starvation, etc.)
There is a potential gain from physical zone isolation, but it protects against a relatively small set of issues. Is it really better to invest in that, or to invest the resources in other safety improvements?
Re: Speed, scale and reliability: 25 years of Google datacenter networking evolution
#67It seems all cutting edge datacenters like x.ai Colossus are using Nvidia networking. Now Google is upgrading to Nvidia networking, too. Since Nvidia owns most of the Gpgpu products, they have top notch networking and interconnect, I wonder if they don't have a plan to own all datacenter hardware in the future. Maybe they plan to also release CPUs, motherboards, storage and whatever else is needed.
https://www.youtube.com/live/Y2F8yisiS6E?si=GbyzzIG8w-mtS7s-...
Re: Speed, scale and reliability: 25 years of Google datacenter networking evolution
#68Earlier quoted context omitted.
Nvidia networking is what used to be called Mellanox networking, which was already dominant in datacenters.
Only within supercomputers (including the smaller GPU ones used to train AI). Normal data centers use Cisco or Juniper or similarly.well known Ethernet equipment, and they still do. The Mellanox/Nvidia Infiniband networks are specifically used for supercomputer-like clusters.
Re: Speed, scale and reliability: 25 years of Google datacenter networking evolution
#69Earlier quoted context omitted.
The phrase "as long as the same overall level of reliability is achieved" is logically flawed when discussing physically co-located vs. geographically separated infrastructure.
Justify that claim. In my experience, the set of issues that would affect 2 buildings close to each other, but not two buildings a mile apart, is vanishingly small, usually just last mile fiber cuts or power issues (which are rare and mitigated by having multiple independent providers), as well as issues like building fires (which are exceedingly rare, we know of, perhaps two of notable impact in more than a decade a…
Re: Speed, scale and reliability: 25 years of Google datacenter networking evolution
#70Earlier quoted context omitted.
Justify that claim. In my experience, the set of issues that would affect 2 buildings close to each other, but not two buildings a mile apart, is vanishingly small, usually just last mile fiber cuts or power issues (which are rare and mitigated by having multiple independent providers), as well as issues like building fires (which are exceedingly rare, we know of, perhaps two of notable impact in more than a decade a…
what happened in gcp paris region then?
It is true, and obvious, that GCP and AWS and Azure use different architectures. It does not obviously follow that any of those architectures are inherently more reliable. And even if it did, it doesn't obviously follow that any of the platforms are inherently more reliable due to a specific architectural decision.
Like, all cloud providers still have regional outages.