Live data from Hacker News

Speed, scale and reliability: 25 years of Google datacenter networking evolution

cloud.google.com

71–80 of 91 posts

Re: Speed, scale and reliability: 25 years of Google datacenter networking evolution

#71
post #64

Earlier quoted context omitted.

The phrase "as long as the same overall level of reliability is achieved" is logically flawed when discussing physically co-located vs. geographically separated infrastructure.

Justify that claim. In my experience, the set of issues that would affect 2 buildings close to each other, but not two buildings a mile apart, is vanishingly small, usually just last mile fiber cuts or power issues (which are rare and mitigated by having multiple independent providers), as well as issues like building fires (which are exceedingly rare, we know of, perhaps two of notable impact in more than a decade a…

I think you're undermining the seriousness of a physical event like a fire. Even if the likelihood of these things is "vanishingly small", the impact is so large that it more than offsets it. Taking the OVH data center fire as an example, multiple companies completely lost their data and are effectively dead now. When you're talking about a company-ending-event, many people would consider even just two examples per decade as a completely unacceptable failure rate. And it's more than just fires: we're also talking about tornados, floods, hurricanes, terrorist attacks, etc.

Google even recognizes this, and suggests that for disaster recovery planning, you should use multiple regions. AWS on the other hand does acknowledge some use cases for multiple regions (mostly performance or data sovereignty), but maintains the stance that if your only concern is DR, then a single region should be enough for the vast majority of workloads.

There's more to the story though, of course. GCP makes it easier to use multiple regions, including things like dual-region storage buckets, or just making more regions available for use. For example GCP has ~3 times as many regions in the US as AWS does (although each region is comparatively smaller). I'm not sure if there's consensus on which is the "right" way to do it. They both have pros and cons.

Re: Speed, scale and reliability: 25 years of Google datacenter networking evolution

#72

Earlier quoted context omitted.

[flagged]

I'm not trying to make you divulge anything. I don't particularly care who you talk to, or who you are, nor do I care if you take it as a "personal insult" that you might be wrong. You are right that it would be nuts that multiple senior people would collude to lie to you, which is why it's almost certainly more likely that you are just misunderstanding the information that was provided to you. It's possible to prove…

You know it is very much region dependent.

You are correct many facilities are owned by the hyperscalers, and they also extensively use colos for hosting entire regions (not only PoPs), specially outside the US. More recently I’d also include Ireland.

I have worked at two cloud providers very close to the netops teams due to my customers, but I have signed NDAs so I won’t go further into it, specially since one of my ex-employers is very touchy about this subject.

Re: Speed, scale and reliability: 25 years of Google datacenter networking evolution

#73

Earlier quoted context omitted.

Yeah there’s a bit of industry worry about that very eventuality — hence the ultra Ethernet consortium trying to work on open source alternatives to the mellanox/nvidia lock-in. https://ultraethernet.org/

Interesting Nvidia is on the steering committee

Cisco have sat on the steering committees for a lot of things where they had a proprietary initial version of something. It's not that unusual, and also, it's often frankly not actually that open; e.g., see the rent seeking racket for access to PCI documentation, or USB-IF actively seeking to prevent open source hardware from existing, etc.

Re: Speed, scale and reliability: 25 years of Google datacenter networking evolution

#74
post #73

Earlier quoted context omitted.

Interesting Nvidia is on the steering committee

Cisco have sat on the steering committees for a lot of things where they had a proprietary initial version of something. It's not that unusual, and also, it's often frankly not actually that open; e.g., see the rent seeking racket for access to PCI documentation, or USB-IF actively seeking to prevent open source hardware from existing, etc.

Eh, the UEC effort is a standards org through the Linux Foundation so it won't be subject to any of the usual chicanery. And actually, it looks like Nvidia is jsut a general member and not one of the Steering Committee members;

https://ultraethernet.org/wp-content/uploads/sites/20/2023/0...

Re: Speed, scale and reliability: 25 years of Google datacenter networking evolution

#75

Speed, scale and reliability Choose any two.

In any decision making matrix you need a constraint that get consumed (economics, size, etc) to force a "choose any two" type situation.

You absolutely can have speed, scale and reliability. You can't have speed, scale, reliability and low cost.

Re: Speed, scale and reliability: 25 years of Google datacenter networking evolution

#76
post #25

Awesome Google... Now learn what an availability zone is and stop creating them with firewalls across the same data center. Oh and make your data centers smaller. Not so big they can be seen in Google Maps. Because otherwise, you will be unable to move those whale sized workloads to an alternative. https://youtu.be/mDNHK-SzXEM?t=564 https://news.ycombinator.com/item?id=35713001 "Unmasking Google Cloud: How to Determi…

To address the availability point of your comment, Google's terminology is slightly different to AWS.

On GCP it sounds like you want to have a multi region architecture, not multi-zone (if you want firewalls outside the same data center).

> Resources that live in a zone, such as virtual machine instances or zonal persistent disks, are referred to as zonal resources. Other resources, like static external IP addresses, are regional. Regional resources can be used by any resource in that region, regardless of zone, while zonal resources can only be used by other resources in the same zone.

https://cloud.google.com/compute/docs/regions-zones

(No affiliation with Google, just had a similar confusion at one point)

Re: Speed, scale and reliability: 25 years of Google datacenter networking evolution

#77
post #62

Earlier quoted context omitted.

To reinforce your point: The scale of cloud data centres reflects the scale of their customer base, not the size of the basket for each individual customer. Larger data centres actually improve availability through several mechanisms: more power components such as generators means the failure of any one is just a few percent instead of a total blackout. You can also partition core infrastructure like routers and powe…

I provided three different references. Despite the massive downvotes on my comment I guess by Google engineers, as a troll...:-)I take comfort on the fact nobody was able to advance a reference to prove me wrong.

AWS Zone is sort-roughly-kinda a GCP Region. It sounds like you want multi-region: https://cloud.google.com/compute/docs/regions-zones

Re: Speed, scale and reliability: 25 years of Google datacenter networking evolution

#78
post #76
post #25

Awesome Google... Now learn what an availability zone is and stop creating them with firewalls across the same data center. Oh and make your data centers smaller. Not so big they can be seen in Google Maps. Because otherwise, you will be unable to move those whale sized workloads to an alternative. https://youtu.be/mDNHK-SzXEM?t=564 https://news.ycombinator.com/item?id=35713001 "Unmasking Google Cloud: How to Determi…

To address the availability point of your comment, Google's terminology is slightly different to AWS. On GCP it sounds like you want to have a multi region architecture, not multi-zone (if you want firewalls outside the same data center). > Resources that live in a zone, such as virtual machine instances or zonal persistent disks, are referred to as zonal resources. Other resources, like static external IP addresses,…

You also need to go multi-region with AWS. I liked their AZ story but in practice it hasn't avoided multi-zone outages (maybe deploys?)

Re: Speed, scale and reliability: 25 years of Google datacenter networking evolution

#79

Earlier quoted context omitted.

[flagged]

This isn't even close to true. You can just go on Google Maps and visually see the literally *hundreds* of wholly-owned and custom built data centers from AWS, MS, and Google. Edge locations (like Cloud CDN) are often in colos, but the main regions compute/storage are not. Most of them are even labeled on Google Maps. Here's a couple search terms you can just type into Google Maps and see a small fraction of what I m…

While the OP is more wrong than right they aren't completely incorrect.

I'm in Australia.

GCP has 2 regions in Australia, Sydney and Melbourne. The Sydney region is in the Equinox DC. Not sure where the Melbourne one is but it isn't a Google-owned facility.

You can see this by comparing Google's Data Center list: https://www.google.com/about/datacenters/locations/ vs their Cloud Location list https://cloud.google.com/about/locations#asia-pacific

Note that the Cloud Locations aren't just "edge": they offer hosting, GPUs etc etc at these locations.

Re: Speed, scale and reliability: 25 years of Google datacenter networking evolution

#80

Earlier quoted context omitted.

[flagged]

Yea, you're not the only "insider" here. And you're 100% wrong. Just because you completely misunderstand what those Amazon/MS employees are doing in those buildings doesn't mean that you know what you're talking about. The big cloud players have the vast majority of their compute and storage hosted out of their own custom built and self-owned data centers. The stuff you see in colos is just the edge locations like C…

This doesn't seem right for GCP.

Compare https://cloud.google.com/about/locations vs https://www.google.com/about/datacenters/locations/

The Cloud locations aren't just edge locations (scroll down on that page and note most have all APIs supported) and there are a lot more of them than there are Google-owned DCs.

Post reply on HN