Live data from Hacker News

Google Cloud Europe service disruption

status.cloud.google.com

61–70 of 153 posts

Re: Google Cloud Europe service disruption

#61
post #56

Title is incorrect, this is not a general outage. There are two separate issues: europe-west-9 (Paris) has been physically flooded with water somehow and is hard down. This is obviously bad if you're using the region in question, but has zero impact elsewhere. https://status.cloud.google.com/incidents/dS9ps52MUnxQfyDGPf... There is a separate issue stopping changes to HTTP load balancers across most of GCP, but it ha…

So this is probably too soon, thoughts and prayers for the datacenter operators and staff out there, but are they going to auction off the flooded hardware? Trying to restore a flooded Google rack sounds like a super fun project. Anyone experience with losing an entire DC to flooding? edit: I just Googled it (lol) and this DC has to be brand spanking new ( https://cloud.google.com/blog/products/infrastructure/google.…

> but are they going to auction off the flooded hardware?

I wonder how many inches/feet we're talking here? The hardware on the top (unless it experienced electrical short) is most likely fine?

Re: Google Cloud Europe service disruption

#62

Earlier quoted context omitted.

that’s how GCP does zones, firewalled off with separate networks/power in the same physical location.

Ouch. Isn't part of separate zones being protected against something, say, like a terrorist attack or a natural disaster that can take down a whole datacenter?

From https://cloud.google.com/docs/geography-and-regions#regions_...

> Regions are independent geographic areas that consist of zones. Zones and regions are logical abstractions of underlying physical resources provided in one or more physical data centers. > (...) > A zone is a deployment area for Google Cloud resources within a region. Zones should be considered a single failure domain within a region. To deploy fault-tolerant applications with high availability and help protect against unexpected failures, deploy your applications across multiple zones in a region.

You should use "region" and "zone" as abstract concepts with shared properties like network topology, local peering, costs, and availability. AFAIK no cloud provider discusses (nor provides guarantees) against specific threats or correlated failures.

There is no guarantee that a given risk will not impact multiple zones, but this risk is lowered by the implementation of various safeguards (for example, rollouts are not happening in multiple regions at the same time).

Google doesn't say "put your VMs in more than one zone because you can be sure we won't have all zones in a region down at the same time", but rather "by putting your VMs in multiple zones in the same region, you can target better SLOs that the SLOs in one zone".

Note that it's different from the concept of "availability zone" of AWS which explicitly says that AZs are physically separated:

> AZs are physically separated by a meaningful distance, many kilometers, from any other AZ, although all are within 100 km (60 miles) of each other.

https://aws.amazon.com/about-aws/global-infrastructure/regio...

Re: Google Cloud Europe service disruption

#63
post #19

Title is incorrect, this is not a general outage. There are two separate issues: europe-west-9 (Paris) has been physically flooded with water somehow and is hard down. This is obviously bad if you're using the region in question, but has zero impact elsewhere. https://status.cloud.google.com/incidents/dS9ps52MUnxQfyDGPf... There is a separate issue stopping changes to HTTP load balancers across most of GCP, but it ha…

I'm not sure if it's a separate issue but I've had trouble creating new VM instances in Google Cloud Console or listing GPU types using their CLI and I'm in europe-west-2. The ticket I was following originally got merged with the Paris flood ticket (by Google). It was working until midnight (London) last night but went down before 8am before recovering about 1h ago for me. Not sure why an outage at one regional data…

Same – was unable to create new VMs in all regions between 7:15am and 11:41am UK time. Not limited to France.

Re: Google Cloud Europe service disruption

#64
post #43
post #16

Earlier quoted context omitted.

It happened at GlobalSwitch Clichy, near Paris. From what I gathered from a french forum[1], it started with a flood and then a fire. No rooms have been affected, apparently. [1]: https://lafibre.info/datacenter/incendie-maitrise-globalswit...

I'm getting horrible flashbacks of OVH DC's those many years ago.

What a disaster. A datacenter made out of wood, what could go wrong ...

Re: Google Cloud Europe service disruption

#66

I thought this title meant cancelled I literally felt the blood leave my face

This was my first thought too. Shows how Google has trained us to expect the worst from them...

TBH, GCP isn't operating like Google does.

I'm a long time customer and have only good things to tell so far.

Re: Google Cloud Europe service disruption

#67
post #43

Earlier quoted context omitted.

I'm getting horrible flashbacks of OVH DC's those many years ago.

THAT was many years ago? Felt like yesterday

A bit more than 2 years: https://www.datacenterdynamics.com/en/news/fire-destroys-ovh...

I thought it was less than 12 months...

Re: Google Cloud Europe service disruption

#68
post #51

Earlier quoted context omitted.

that’s how GCP does zones, firewalled off with separate networks/power in the same physical location.

Are you joking? Please tell me that’s a joke, because there’s no way a cloud provider that big could be that daft. If that’s true, what’s the fucking point of separating them at all?

Because power / network / software maintenance events cause outages. Those are scheduled per zone, and so they will take down one zone but not a whole data center.
Post reply on HN