Live data from Hacker News

Google Cloud Europe service disruption

status.cloud.google.com

141–150 of 153 posts

Re: Google Cloud Europe service disruption

#141

Title is incorrect, this is not a general outage. There are two separate issues: europe-west-9 (Paris) has been physically flooded with water somehow and is hard down. This is obviously bad if you're using the region in question, but has zero impact elsewhere. https://status.cloud.google.com/incidents/dS9ps52MUnxQfyDGPf... There is a separate issue stopping changes to HTTP load balancers across most of GCP, but it ha…

Per [1], there was a related issue affecting Cloud Console operations globally, starting from the point where the incident went regional at 23:00 PDT, and lasting until 02:00 PDT-ish. It is incorrect to say that this had zero impact elsewhere.

Sounds like some global control plane related to instance management operations started returning errors once one region failed. Or perhaps it was just the UI frontend?

[1] https://status.cloud.google.com/incidents/BWK7QzFBmfaZ4iztke...

Re: Google Cloud Europe service disruption

#142
post #56

Earlier quoted context omitted.

So this is probably too soon, thoughts and prayers for the datacenter operators and staff out there, but are they going to auction off the flooded hardware? Trying to restore a flooded Google rack sounds like a super fun project. Anyone experience with losing an entire DC to flooding? edit: I just Googled it (lol) and this DC has to be brand spanking new ( https://cloud.google.com/blog/products/infrastructure/google.…

I once was a customer of a DC who's roof drainage was clogged, turning it into a lake after a couple of rain storms. It then proceeded to rain inside the DC as the roof started to leak from all the pressure. "Servers are down, I'll head over to the DC" turned into "Um... it's raining _in the DC_. Get me some tarps and get us cut over to the backup in the office". Ah, the glory days of running out of a single co-lo ac…

Many years ago, I managed a server room with dedicated cooling on the 4th floor of a 4-story building with a flat roof. One night the temp alarms went off, and when I showed up water was dripping off my overhead Liebert unit and onto the racks.

And it wasn't even raining outside! So I grab some plastic to cover the racks and phone in emergency portable cooling as the room's AC started failing.

It turns out earlier that day, a technician performing seasonal maintenance on a boiler tank on the roof had drained the tank and refilled it. But instead of directing the water out into a proper drain, he sent it down a convenient pipe that was actually a vent from our ceiling into the boiler house. The boiler was dozens of meters from my server room, but the water followed the old steel and plaster ceiling remnants over to my computers.

And this boiler water was more exciting than rain: it came with all the dissolved minerals, metals, and preservatives computers crave! I didn't lose any computers in the racks, but it killed the Liebert's control board.

Re: Google Cloud Europe service disruption

#143
post #126

Earlier quoted context omitted.

GCP has multiple zones in the same physical building. Not all cloud providers have distinct physical buildings for each Availability Zone.

Do they have an official description what a zone is somewhere? Back in the days when we had our own data centers a zone was defined as a "fire section" meaning that it should not be impacted if any other zone of the data center had a fire. This obviously means that you can't call 3 floors of a building a zone. Edit: The information on this site https://cloud.google.com/docs/geography-and-regions#regions_... clearly s…

Physically distinct could refer to distinct hardware in the same building and cage space. It’s “physically distinct”. Google makes no promises that the zones are in different buildings or separated by N feet/miles of space.

Re: Google Cloud Europe service disruption

#144
post #56

Earlier quoted context omitted.

So this is probably too soon, thoughts and prayers for the datacenter operators and staff out there, but are they going to auction off the flooded hardware? Trying to restore a flooded Google rack sounds like a super fun project. Anyone experience with losing an entire DC to flooding? edit: I just Googled it (lol) and this DC has to be brand spanking new ( https://cloud.google.com/blog/products/infrastructure/google.…

The machines are not industry standard stuff, and they don't auction, they destroy for customer security. See here: https://www.datacenterknowledge.com/google-alphabet/robots-n...

Just the drives are destroyed. The servers themselves end up in all sorts of spots:

https://www.ebay.com/b/Google-Server/11211/bn_7023306662

Re: Google Cloud Europe service disruption

#145
post #124

Earlier quoted context omitted.

2015 Chennai (South India) Floods. It was the flood of a century. [1] Our DC was intact, but the building and access was cut-off. We lost the backup diesel power generators in the flooding. Of course, grid power was cut-off. Our DC operating team managed to shutdown all the servers and racks cleanly before UPS power was completely drained. The 4 engineers and 2 security guards then swam out of the compound in chest h…

Ha, you bring back old memories. We had the largest compute footprint in India at that time in Ambattur (Chennai industrial suburb). This particular DC in question was as multi-story building and the ground-floor itself was several ft above road level and there was the huge natural lake in front. Luckily heavy rains only caused havoc to the road-side storm-drains and road traffic. And we had more than 250K liters of…

Yes. It was a miracle that Ambattur did not suffer as much given the proximity to Redhills lake reservoir. Had the Water Resource Department also opened the sluice gates of the Redhills reservoir like Chembarampakkam lake during the floods and incessant rains, the situation would have been different. Given Ambattur was accessible and relatively unaffected, that was the location we brought up our alternate operating site within a week.

In any case, it is good you didn't have to go through a DC recovery during one of the worst disasters in the 21st century.

The question I keep asking in all DR planning sessions/table top exercises is - what would we do if we had a situation like what happened in Fukushima or in Chennai 2015. In both cases, flooding caused failure of backup power generators. Also, what do we do when we have all or partial resources, but are faced with a denial-of-premises situation (what I faced).

Re: Google Cloud Europe service disruption

#146

Earlier quoted context omitted.

The machines are not industry standard stuff, and they don't auction, they destroy for customer security. See here: https://www.datacenterknowledge.com/google-alphabet/robots-n...

Just the drives are destroyed. The servers themselves end up in all sorts of spots: https://www.ebay.com/b/Google-Server/11211/bn_7023306662

Those are all Google search appliances, Google sold those. They're not operated by Google themselves.

Re: Google Cloud Europe service disruption

#147

Title is incorrect, this is not a general outage. There are two separate issues: europe-west-9 (Paris) has been physically flooded with water somehow and is hard down. This is obviously bad if you're using the region in question, but has zero impact elsewhere. https://status.cloud.google.com/incidents/dS9ps52MUnxQfyDGPf... There is a separate issue stopping changes to HTTP load balancers across most of GCP, but it ha…

"europe-west-9 (Paris) has been physically flooded [...], but has zero impact elsewhere."

I am afraid this is not true. We have nothing in europe-west-9, but problem in this region caused global problem with Cloud Console, which hit us, because we were not able to use it for several hours.

Snippert from https://status.cloud.google.com/incidents/dS9ps52MUnxQfyDGPf...:

"Cloud Console: Experienced a global outage, which has been mitigated. Management tasks should be operational again for operations outside the affected region (europe-west9). Primary impact was observed from 2023-04-25 23:15:30 PDT to 2023-04-26 03:38:40 PDT."

Re: Google Cloud Europe service disruption

#148

Earlier quoted context omitted.

But does it really matter that the incident is a flood or a cascading software failure if the likelihood and severity is the same? Being in the same building is an "implementation detail" from a customer perspective, what matters is the consequences of this decision. For example, maybe this decision allows for better network connectivity at a lower cost for inter-zones traffic, while, on the other hand, not protectin…

It absolutely does matter. The MTTR for outages caused by physical damage is way higher, and resiliency against physical disasters is a major selling point of availability zones as a fault container. Hosting every zone of your region (if that's actually the case here) in the same building is simply negligent. Besides the obvious risks like this incident, even if the zones have physical fire barriers, chances that ope…

True, I implicitly included the MTTR in the "severity", but this is actually a different thing (severity is more about the impact radius).

But I don't think it changes my point: knowing what/how Google Cloud designs regions or zones is still an implementation detail, what matters is what MTTR they are targeting and this should be known ahead of time.

There are so many "implementation details" that customers are not aware of, because they are always changing, non contractual, or just hard to make sense of, what matters is meaningful abstractions.

I am not saying it's OK if the zones are in the same building or not, I don't know and I was really surprised when I discovered this a few years ago. But this information gives you a mental model of "what could go wrong" that is biased towards some specific risks, and in my experience, relying on these very practical aspects make the risk analysis and design decisions harder to make.

Otho, one thing that may be problematic too (and biasing) is that the common understood definition of a "zone" is the one people know from AWS, so using the same term without being very explicit about the differences will also lead to incorrectly calculated risks. I find the public documentation of Google Cloud too vague in general (and often ambiguous).

Re: Google Cloud Europe service disruption

#149
post #115

Earlier quoted context omitted.

Ah, interesting; it's been a while since I played with AWS, and that service wasn't there back then. I'm guessing that allocating a new AWS Global Accelerator address takes a while?

I've only done it once (the way they have it architected, it's a "set and forget" sort of thing, your LB changes don't touch the Global Accelerator) but I do seem to recall that it took awhile to create the resource. Maybe 5-10 minutes?

5-10 minutes is accurate for creating and rolling out changes to a Global Accelerator

Re: Google Cloud Europe service disruption

#150
post #126

Earlier quoted context omitted.

GCP has multiple zones in the same physical building. Not all cloud providers have distinct physical buildings for each Availability Zone.

Do they have an official description what a zone is somewhere? Back in the days when we had our own data centers a zone was defined as a "fire section" meaning that it should not be impacted if any other zone of the data center had a fire. This obviously means that you can't call 3 floors of a building a zone. Edit: The information on this site https://cloud.google.com/docs/geography-and-regions#regions_... clearly s…

I could not find the GCP equivalent to this from AWS:

"AZs are physically separated by a meaningful distance, many kilometers, from any other AZ, although all are within 100 km (60 miles) of each other."

https://aws.amazon.com/about-aws/global-infrastructure/regio...

Post reply on HN