Live data from Hacker News

Google Cloud Europe service disruption

status.cloud.google.com

131–140 of 153 posts

Re: Google Cloud Europe service disruption

#131
post #16

Earlier quoted context omitted.

It happened at GlobalSwitch Clichy, near Paris. From what I gathered from a french forum[1], it started with a flood and then a fire. No rooms have been affected, apparently. [1]: https://lafibre.info/datacenter/incendie-maitrise-globalswit...

If it's the one in Clichy I'm thinking of it's dug into the embankment that lines a railway basin, so... yeah, floods suck.

It is the Clichy's one. It's not that dug into, where did you get that from? (Used to work there circa 2010). I think the water retention made its way to the battery rooms. No recent floods (nor rain) in Paris (nor most of france) lately.

Re: Google Cloud Europe service disruption

#132

Earlier quoted context omitted.

AWS has similarly suffered outages from an entire datacenter being taken out like this. No one is immune. If you want true fault-tolerance you need to be multi-regional (everyone says as much), ideally, multi-continental. europe-west9 is the only large Google datacenter in France afaik. Building more would cost lots more money, and it seems like the market isn't there for it. Workloads that require data locality in F…

Eh source for that. AWS has had issues where a single Zone caused such a lack of capacity in the region that some multi-zone services degraded to the point of a domino fail-over. However I've not heard of any AWS event where a fire/flood in AZ A also caused a fire/flood in AZ B.

It sounds like this might just be confusion over nomenclature, with Google and Amazon using different terms for the same thing.

Regardless, with GCP, if you need redundancy that can survive the loss of an entire datacenter, then you need to be multi-regional. This has been widely known best practice for a long time.

Re: Google Cloud Europe service disruption

#133
post #124
post #56

Earlier quoted context omitted.

So this is probably too soon, thoughts and prayers for the datacenter operators and staff out there, but are they going to auction off the flooded hardware? Trying to restore a flooded Google rack sounds like a super fun project. Anyone experience with losing an entire DC to flooding? edit: I just Googled it (lol) and this DC has to be brand spanking new ( https://cloud.google.com/blog/products/infrastructure/google.…

2015 Chennai (South India) Floods. It was the flood of a century. [1] Our DC was intact, but the building and access was cut-off. We lost the backup diesel power generators in the flooding. Of course, grid power was cut-off. Our DC operating team managed to shutdown all the servers and racks cleanly before UPS power was completely drained. The 4 engineers and 2 security guards then swam out of the compound in chest h…

Ha, you bring back old memories. We had the largest compute footprint in India at that time in Ambattur (Chennai industrial suburb). This particular DC in question was as multi-story building and the ground-floor itself was several ft above road level and there was the huge natural lake in front. Luckily heavy rains only caused havoc to the road-side storm-drains and road traffic. And we had more than 250K liters of diesel to last us more than 24 hours and we had several tankers on standby. So we didn't have to shutdown anything. Funny thing is we had selected this site less than a year ago and had discussed the 100 year flood lines and worst case probabilities of heavy rains and flooding etc. Being well-prepared really paid off.

Re: Google Cloud Europe service disruption

#134
post #45
post #19

Earlier quoted context omitted.

I'm not sure if it's a separate issue but I've had trouble creating new VM instances in Google Cloud Console or listing GPU types using their CLI and I'm in europe-west-2. The ticket I was following originally got merged with the Paris flood ticket (by Google). It was working until midnight (London) last night but went down before 8am before recovering about 1h ago for me. Not sure why an outage at one regional data…

Cloud Console is having issues related to the outage in europe-west9 > Customer using Cloud Console globally are unable to open and view the Compute Engine related pages like: Instance creation page Disk creation page Instance templates page Instance Groups page https://status.cloud.google.com/incidents/dS9ps52MUnxQfyDGPf...

I got errors trying to open the instance group list and we don't have any resources in europe-west9.

Re: Google Cloud Europe service disruption

#135
post #51

Earlier quoted context omitted.

that’s how GCP does zones, firewalled off with separate networks/power in the same physical location.

Are you joking? Please tell me that’s a joke, because there’s no way a cloud provider that big could be that daft. If that’s true, what’s the fucking point of separating them at all?

[deleted]

Re: Google Cloud Europe service disruption

#136
post #73
post #51

Earlier quoted context omitted.

Are you joking? Please tell me that’s a joke, because there’s no way a cloud provider that big could be that daft. If that’s true, what’s the fucking point of separating them at all?

Minimising the blast radius from logical changes (software & config) that get rolled out at an AZ-level. Their descriptions[0] however promise zones have a "high degree of independence from one another in terms of physical and logical infrastructure". Just how well separated this physical zonal infrastructure was remains to be seen ... [0] https://cloud.google.com/architecture/disaster-recovery#regi...

Yeah I feel like that description is a lie. Some customers would probably think twice about putting things into the same region if they knew zones weren't physically separated, or go to AWS.

Re: Google Cloud Europe service disruption

#137

Earlier quoted context omitted.

Eh source for that. AWS has had issues where a single Zone caused such a lack of capacity in the region that some multi-zone services degraded to the point of a domino fail-over. However I've not heard of any AWS event where a fire/flood in AZ A also caused a fire/flood in AZ B.

But does it really matter that the incident is a flood or a cascading software failure if the likelihood and severity is the same? Being in the same building is an "implementation detail" from a customer perspective, what matters is the consequences of this decision. For example, maybe this decision allows for better network connectivity at a lower cost for inter-zones traffic, while, on the other hand, not protectin…

Seems the likelihood isn't the same. AWS is separating AZs physically, GCP is not. I'd want to know this as a customer, not some abstraction.

Re: Google Cloud Europe service disruption

#138
post #85
post #76

Earlier quoted context omitted.

I'm not sure what the disk encryption story is in Google Cloud but I'd rather it didn't end up on Ebay. Mind you, "flooded" covers a wide range of possibilities and a surprisingly small amount of water ingress would trip a breaker while leaving the racks in good order.

> a surprisingly small amount of water ingress would trip a breaker while leaving the racks in good order. If that were the case they wouldn't be saying "There is no current ETA for recovery," and "it is expected to be an extended outage. Customers are advised to failover to other regions."

Having (for example) 6 inches of water in your 115kV switch room is a small-scale problem that can cause a large-scale outage.

Re: Google Cloud Europe service disruption

#139

Title is incorrect, this is not a general outage. There are two separate issues: europe-west-9 (Paris) has been physically flooded with water somehow and is hard down. This is obviously bad if you're using the region in question, but has zero impact elsewhere. https://status.cloud.google.com/incidents/dS9ps52MUnxQfyDGPf... There is a separate issue stopping changes to HTTP load balancers across most of GCP, but it ha…

Wow, “physically flooded with water somehow” and “load balancers” config propagation issue are so drastically different! Good reminder that downtime happens for many wild reasons, and you may want to take 30 seconds and set up a free website / API monitor with Heii On-Call [1] because we would have alerted you to either of these issues if they affected your app. Really, a simple HTTP probe provides tremendous monitor…

We just use a simple cloud function for that.

Re: Google Cloud Europe service disruption

#140

Earlier quoted context omitted.

Eh source for that. AWS has had issues where a single Zone caused such a lack of capacity in the region that some multi-zone services degraded to the point of a domino fail-over. However I've not heard of any AWS event where a fire/flood in AZ A also caused a fire/flood in AZ B.

But does it really matter that the incident is a flood or a cascading software failure if the likelihood and severity is the same? Being in the same building is an "implementation detail" from a customer perspective, what matters is the consequences of this decision. For example, maybe this decision allows for better network connectivity at a lower cost for inter-zones traffic, while, on the other hand, not protectin…

It absolutely does matter.

The MTTR for outages caused by physical damage is way higher, and resiliency against physical disasters is a major selling point of availability zones as a fault container.

Hosting every zone of your region (if that's actually the case here) in the same building is simply negligent.

Besides the obvious risks like this incident, even if the zones have physical fire barriers, chances that operators will be allowed in to one "zone" after another has a fire are slim to none.

Post reply on HN