Live data from Hacker News

Speed, scale and reliability: 25 years of Google datacenter networking evolution

cloud.google.com

31–40 of 91 posts

Re: Speed, scale and reliability: 25 years of Google datacenter networking evolution

#31
They managed to double from 6 Petabit per second in 2022 to 13 Pbps in 2023. I assume with ConnectX-8 this could be 26 Pbps in 2025/26. The ConnextX-8 is PCI-e 6 so I assume we could get 1.6Tbps ConnextX-9 with PCI-e 7.0 which is not far away.

Cant wait to see the FreeBSD Netflix version of that post.

This also goes back to how increasing throughput is relatively easy and has a very strong roadmap. While increasing storage is difficult. I notice YouTube has been serving higher bitrate video in recent years with H.264. Instead of storing yet another copy of video files in VP9 or AV1 unless they are 2K+.

Re: Speed, scale and reliability: 25 years of Google datacenter networking evolution

#32
post #30

Earlier quoted context omitted.

Nvidia networking is what used to be called Mellanox networking, which was already dominant in datacenters.

Only within supercomputers (including the smaller GPU ones used to train AI). Normal data centers use Cisco or Juniper or similarly.well known Ethernet equipment, and they still do. The Mellanox/Nvidia Infiniband networks are specifically used for supercomputer-like clusters.

Mellanox Ethernet NIC got used a bunch of places due to better programmability.

Re: Speed, scale and reliability: 25 years of Google datacenter networking evolution

#33

It seems all cutting edge datacenters like x.ai Colossus are using Nvidia networking. Now Google is upgrading to Nvidia networking, too. Since Nvidia owns most of the Gpgpu products, they have top notch networking and interconnect, I wonder if they don't have a plan to own all datacenter hardware in the future. Maybe they plan to also release CPUs, motherboards, storage and whatever else is needed.

Yeah there’s a bit of industry worry about that very eventuality — hence the ultra Ethernet consortium trying to work on open source alternatives to the mellanox/nvidia lock-in.

https://ultraethernet.org/

Re: Speed, scale and reliability: 25 years of Google datacenter networking evolution

#34

Does gcp have the worst networking for gpu training though?

For TPU pods they use 3D torus topology with multi-terabit cross connects. For GPU, A3 Ultra instances offer "non-blocking 3.2 Tbps per server of GPU-to-GPU traffic over RoCE".

Is that the worst for training? Namely: do superior solutions exist?

Re: Speed, scale and reliability: 25 years of Google datacenter networking evolution

#35
post #17

This mentions Jupiter generations, which I think is about 10-15 years old at this point. It doesn't really talk about what existed before so it's not really 25 years of history here. I want to say "Watchtower" was before Jupiter? but honestly it's been about a decade since I read anything about it. Google's DC networking is interesting because of how deeply integrated it is into the entire software stack. Click on so…

Nvidia got ConnectX from their Mellanox acquisition -- they were experts in RMDA, particularly with Infiniband but eventually pushing Ethernet (RoCE). These NICs have hardware-acceleration of RDMA. Over the RDMA fabric, GPUs can communicate with each other without much CPU usage (the "GPU-to-GPU" mentioned in the article).

[I know nothing about Jupiter, and little about RDMA in practice, but used ConnectX for VMA, its hardware-accelerated, kernel-bypass tech.]

Re: Speed, scale and reliability: 25 years of Google datacenter networking evolution

#36
post #17

This mentions Jupiter generations, which I think is about 10-15 years old at this point. It doesn't really talk about what existed before so it's not really 25 years of history here. I want to say "Watchtower" was before Jupiter? but honestly it's been about a decade since I read anything about it. Google's DC networking is interesting because of how deeply integrated it is into the entire software stack. Click on so…

From memory: Firehose > Watchtower > WCC > SCC > Jupiter v1

Re: Speed, scale and reliability: 25 years of Google datacenter networking evolution

#37
post #25

Awesome Google... Now learn what an availability zone is and stop creating them with firewalls across the same data center. Oh and make your data centers smaller. Not so big they can be seen in Google Maps. Because otherwise, you will be unable to move those whale sized workloads to an alternative. https://youtu.be/mDNHK-SzXEM?t=564 https://news.ycombinator.com/item?id=35713001 "Unmasking Google Cloud: How to Determi…

[flagged]

Re: Speed, scale and reliability: 25 years of Google datacenter networking evolution

#38
post #25

Awesome Google... Now learn what an availability zone is and stop creating them with firewalls across the same data center. Oh and make your data centers smaller. Not so big they can be seen in Google Maps. Because otherwise, you will be unable to move those whale sized workloads to an alternative. https://youtu.be/mDNHK-SzXEM?t=564 https://news.ycombinator.com/item?id=35713001 "Unmasking Google Cloud: How to Determi…

[flagged]

This isn't even close to true. You can just go on Google Maps and visually see the literally *hundreds* of wholly-owned and custom built data centers from AWS, MS, and Google. Edge locations (like Cloud CDN) are often in colos, but the main regions compute/storage are not. Most of them are even labeled on Google Maps.

Here's a couple search terms you can just type into Google Maps and see a small fraction of what I mean:

- "Google Data Center Berkeley County"

- "Microsoft Data Center Boydton"

- "GXO council bluffs" (two locations will appear, both are GCP data centers)

- "Google Data Center - Henderson"

- "Microsoft - DB5 Datacentre" (this one is in Dublin, and is huuuuuge)

- "Meta Datacenter Clonee"

- "Google Data Center (New Albany)" (just to the east of this one is a massive Meta data center campus, and to the immediate east of it is a Microsoft data center campus under construction)

And that's just a small sample. There are hundreds of these sites across the US. You're somewhat right that a lot of international locations are colocated in places like Equinix data centers, but even then it's not all of them and varies by country (for example in Dublin they mostly all have their own buildings, not colo). If you know where to look and what the buildings look like, the custom-build and self-owned data centers from the big cloud providers are easy to spot since they all have their own custom design.

Re: Speed, scale and reliability: 25 years of Google datacenter networking evolution

#39

Earlier quoted context omitted.

[flagged]

This isn't even close to true. You can just go on Google Maps and visually see the literally *hundreds* of wholly-owned and custom built data centers from AWS, MS, and Google. Edge locations (like Cloud CDN) are often in colos, but the main regions compute/storage are not. Most of them are even labeled on Google Maps. Here's a couple search terms you can just type into Google Maps and see a small fraction of what I m…

[flagged]

Re: Speed, scale and reliability: 25 years of Google datacenter networking evolution

#40

Earlier quoted context omitted.

This isn't even close to true. You can just go on Google Maps and visually see the literally *hundreds* of wholly-owned and custom built data centers from AWS, MS, and Google. Edge locations (like Cloud CDN) are often in colos, but the main regions compute/storage are not. Most of them are even labeled on Google Maps. Here's a couple search terms you can just type into Google Maps and see a small fraction of what I m…

[flagged]

Yea, you're not the only "insider" here. And you're 100% wrong. Just because you completely misunderstand what those Amazon/MS employees are doing in those buildings doesn't mean that you know what you're talking about.

The big cloud players have the vast majority of their compute and storage hosted out of their own custom built and self-owned data centers. The stuff you see in colos is just the edge locations like Cloudfront and Cloud CDN, or the new-ish offerings like AWS Local Zones (which are a mix between self-owned and colo, depending on how large the local zone is).

Most of this is publicly available by just reading sites like datacenterdynamics.com regularly, btw. No insider knowledge needed.

Post reply on HN