Live data from Hacker News

Google Cloud region currently down due to water intrusion

status.cloud.google.com

141–150 of 187 posts

Re: Google Cloud region currently down due to water intrusion

#141

Not sure what kind of fire there was there, but once those automatic sprinkler systems get going, they are very difficult to stop. Someone in my freshman college dorm decided to use one as a clothes hanger hook and broke the thermometer in there. The sprinkler damaged the entire floor with water and the floor below had spotty rain as well. The fire department came and was mainly concerned about evacuating everyone ra…

datacenters dont use sprinkler systems (or at least they should not).

A non-water fire suppression system for a 300,000+ square feet warehouse-scale datacenter would be incredibly expensive.

Re: Google Cloud region currently down due to water intrusion

#142
post #138

[disclaimer: SRE @ Google, I was involved with the incident, obvious conflicts of interest] Hey Dang, thanks for cleaning up the thread. One thing to note is that the title is not correct. The entire region is not currently down, as the regional impact was mitigated as of 06:39 PDT, per the support dashboard (though I think it was earlier). The impact is currently zonal (europe-west9-a), so having zone in the title a…

Would you be able to comment a bit on the emotional (perhaps there’s a better word) aspect of the response?

Was there a lot of anxiety? Panic? Or was it just a “woof that sucks. Time to follow a checklist and then do a bunch of paper work” ?

What I’m curious about is what it feels like on a team at a company like Google when there is a major system failure.

Re: Google Cloud region currently down due to water intrusion

#143
post #46

I can't ignore the feeling that Google Cloud is sub par compared to AWS. How did this again cause a multi zone failure. Why haven't they fixed those dependencies the last few times they had a full region failure.

> I can't ignore the feeling that Google Cloud is sub par compared to AWS. That goes without saying at this point. More importantly, it’s proven worse than Azure.

Depending on the metric. I have not seen any incidents where the isolation between customers was breach. Azure had several. Their compute offerings are better. We could go on.

On the other hand, Azure was(and still is) upfront about not having AZs - now that they have rolled out, hopefully those are not in the same building.

Re: Google Cloud region currently down due to water intrusion

#144
post #92
post #28

Earlier quoted context omitted.

Yes. Azure and GCPs numbers on the size of their AZs and such are more marketing spin than hard engineering. AWS keeps these in separate physical locations to provide true separation. While there have been tech related regional incidents at AWS a physical event disabling multiple AZs would be extremely unlikely given their much more robust and geographically distributed design. If such a physical event had happened i…

Googler, opinions are my own. I think you misunderstand Google's infrastructure. I'm guessing that each GCP zone is actually a Borg Cell (see: https://storage.googleapis.com/pub-tools-public-publication-... ). Borg cells tend to be isolated from eachother in many ways in the physical layer (networking and management being a big one, not sure about power). So networking or machine management for an entire zone could g…

Yeah, you’re not getting what people are saying. AWS’s AZs are much more separated than GCPs. Your recommendation that one could build across regions isn’t what folks are talking about here since there is a big benefit to having geographically separate AZs in the same region. That’s where GCP is falling short here.

Re: Google Cloud region currently down due to water intrusion

#145

Earlier quoted context omitted.

> Also, why batteries in a datacenter? Everything serious in the telecom/ISP infrastructure sector has a big -48VDC battery plant, or preferably separate A and B side -48VDC battery plants, to provide a significant buffer between power going Grid --> AC-to-DC Rectifiers --> Equipment, and when a generator can start up, warm up, and transfer switch does its job. Even if a bunch of servers don't have any UPS or battery…

Very different trade offs in play for google who run with a relatively high tolerance for failure at the individual machine or even rack level. At one point I believe there were batteries in every rack, though I don’t know what they're building these days. A telco DC is gonna have more network interconnect with lower tolerance for failure due to capacity impact that isn’t easy to double. Think like a fiber terminatio…

What I was saying above is that the 'core' of a google DC has a massive amount of network interconnect and needs for battery backup not very different from a big IX point or traditional "primary CO" for a city in a telco environment.

By square footage maybe 95% of a google DC might have no UPS or battery backup but the core network for things like routers and DWDM transport equipment absolutely will have such.

If they were unlucky enough that the burst cooling loop met with the battery plant for the core gear in a building or small campus of buildings....

Re: Google Cloud region currently down due to water intrusion

#146

Earlier quoted context omitted.

Anyone talking up firefighters like this clearly hasn't been around them much. They're boys with toys that they don't frequently get to use and they work for the government. Follow the incentives. They'll do their jobs but they don't give a lot of fucks about things like "unnecessary property damage" and "other people's financial well being" and anything else not written in their KPIs. I used to drive tow truck. I ca…

Firefighters are personally incentivized to take photos like this: https://nypost.com/2022/01/10/fireman-in-post-photo-recalls-... . They are not incentivized to protect your stuff, that's what insurance does. Good. I want you to save me and my family from dying in a fucking fire. This thread is obsessed with saving Funko Pop collections for some reason??

That's just a textbook appeal to emotion.

Most of their calls are mundane stuff. And they leave a pretty decently wide path of destruction in doing that. We're talking like mundane situations where there is no urgency and no need to tear shit up in the interest of time.

I once arrived to a minor rollover after the cops but before fire. Nobody injured. Occupant trapped because she was a large lady and couldn't release her seatbelt upside down and was having difficulty unlocking the car because side curtain airbags.

I offered to flip the car and treat it like a lockout. "Customer" was fine with it. Cop was iffy. FD showed up, didn't want to hear it, broke the window, unlocked the car, opened the door, cut her belt rather than release it and dropped her on her face and then had difficulty getting her out. Now I get that they have "procedures" but this seems like a forest for the trees situation.

Or they'll show up, shut down two lanes for a minor fire on the shoulder and not move the trucks until the car is loaded on a tow truck and gone. Supposedly it's to keep them safe from being hit by traffic. Meanwhile here I am not blocking traffic to recover shit that broke down.

Sure, they'll save a life if the situation presents itself but they sure don't care about being tidy about it.

Re: Google Cloud region currently down due to water intrusion

#147
post #93

Earlier quoted context omitted.

Plot twist: the server racks were made out of sodium.

You're not far off: the batteries are (probably) made of lithium. Also, why batteries in a datacenter? When you implement a flush() command at the lowest level you're faced with two choices: 1) actually write to disk, then return from the call, 2) write to some cache/RAM and have just enough battery locally to ensure that you can write it to disk even if all power goes out. Then there's the other problem of surviving…

The batteries in the datacenter are simply there to hold the power until the generators are all up and running, and the phases are in sync.

They create 3 separate arrays of batteries in each back. Each array represents a power phase, A-B-C.. if I remember correctly, each array has a number of low voltage/2000 amp batteries connected in series to make up for a 2000amp 480 volt leg on the other end.

In a tier 4 plus+1 datacenter, they have 4 battery rooms and 4 generators for each data pod. You have a primary generator and UPS battery set, and a backup generator set for each pod. And then that generator set has its own primary and secondary backup set. The end result is that they can work on any piece of equipment without interrupting power. In the event they lost the primary set or needed to take it offline for maintenance, they have the whole secondary redundant set to fallback on.

The servers on the received on the power cord after it passes the switchgear never know that there has been power source changes on the other end.

Re: Google Cloud region currently down due to water intrusion

#148
post #127

Earlier quoted context omitted.

Yeah, I always thought datacenters would use Halon. It of course has the problem of suffocating everyone.

Halon has been banned for years because 1) it's bad for the ozone layer and 2) it'll kill you. Newer systems (FM-200, Inergen, etc.) fight the fire by removing heat instead of removing oxygen.

Halon is still used. Unfortunately the same properties that makes it effective also makes it harm the ozone layer. It does not just remove heat or oxygen, it directly interferes with the reaction involved in combustion, making things stop burning.

Re: Google Cloud region currently down due to water intrusion

#149
post #5

Earlier quoted context omitted.

There's an interesting Twitter thread about that topic here: https://twitter.com/GergelyOrosz/status/1651256082424012806 Based on that thread it sounds like only AWS guarantees that their AZs are in physically separate DCs, while for Google and Microsoft AZs could be in separate buildings of the same DC facility.

I would really like to see the physical DC separation at "The Dalles, Oregon".

Looks like there are three buildings[1] to me, not entirely sure what goes where, obviously.

1: https://goo.gl/maps/Tfw5UpSsoYiN3YMVA

Re: Google Cloud region currently down due to water intrusion

#150

I can imagine clients who used one DC being impacted. But Google’s services would be designed for a single DC going down, right? Data would be eventually consistent (once they find and plug the hard drives in) but isn’t this the promise of the cloud and they’re (approximately) the best at using it. I have to assume it’s a fault that not even distributed services can paper over. Eg lots of crucial data in flight and t…

Cloud customers have no control on which or how many 'datacenters' are used. That's not something that's even advertised or easily available to customers.

The logical units are regions and availability zones or the equivalent nomenclature in each cloud. One availability zone is expected to be one or more datacenters.

We have thousands of instances in AWS. I do not know - or care - where they are physically located(other than the region name, say, Oregon). I expect at most one availability zone to get impacted if a datacenter goes up in flames (and sometimes, just a portion of one). I mention in another comment that AWS has had issues before and production systems barely got impacted. And recovered with zero intervention - instances with failed health checks get replaced by brand new ones in whatever AZs are still operational.

> Data would be eventually consistent (once they find and plug the hard drives in)

At the level of abstractions cloud operates, no-one is plugging drives in – someone is, but you can never see it.

Most cloud workloads use network attached storage - when you can even see the logical drives (SaaS offerings may not even have that abstraction). We don't know (or care) how many physical hard drives exist, or where they are. Latency requirements probably dictate that they are close to the actual instances, but there's usually data replication going on even across DCs.

In addition to that, at least in AWS, if you have saved any volume snapshots at all, they will be in S3. This data will be replicated and underlying systems can even use it to restore lost or corrupted data without you even noticing and sometimes even without a recent snapshot, as storage keeps track of what blocks have been rewritten since the last snapshot. In a particularly bad case you might have to do a restore.

In almost a decade and number of volumes in the 6 digits (no clue how many drives that is!) we never had a single volume fail on AWS. Some got into a 'degraded' state and then recovered.

We haven't had any failures on GCP either. In the case of GCP, even faulty hypervisors are transparently worked around - we never notice other than some audit logs saying the VM was moved. They even preserve the network connections. AWS requires a stop/start to do the same, but your VM will be up and running in a different hypervisor (sometimes a different datacenter) in a couple of minutes, with all the storage.

Mind you, AWS promises eleven nines(!) of durability for S3.

When you do have locally attached storage, it's treated as ephemeral and it's gone if the instance restarts.

> I have to assume it’s a fault that not even distributed services can paper over.

If a single datacenter fails, since it _should_ be at most one AZ(this case seems to be different) that will depend on how the application is architected. Requests in flight will obviously fail, how big of a deal depends on the problem domain. For most web apps, this will cause a retry and that's the end of the story, others will be specifically engineered to deal with receiving multiple messages or dropping messages. For example, if you need at most once delivery guarantees, you need to take extra measures

Not all applications can survive an entire region going down. Some can, but that usually raises costs if you are continuously replicating data across regions. If you do that, then you should be able to steer traffic to the surviving regions. You can do that old-school by changing DNS records, or you could have fanciers solutions such as global anycast loadbalancers and have a single IP worldwide that still goes to the closest healthy region.

Post reply on HN