Live data from Hacker News

Google Cloud Europe service disruption

status.cloud.google.com

151–153 of 153 posts

Re: Google Cloud Europe service disruption

#151
post #53

I thought this title meant cancelled I literally felt the blood leave my face

It was probably striking for better working conditions so Google terminated it

A memorable observation from idlewords is that Googlers will organize for better conditions for Googlers but they never organize for better conditions for their users.

Re: Google Cloud Europe service disruption

#152

Earlier quoted context omitted.

It absolutely does matter. The MTTR for outages caused by physical damage is way higher, and resiliency against physical disasters is a major selling point of availability zones as a fault container. Hosting every zone of your region (if that's actually the case here) in the same building is simply negligent. Besides the obvious risks like this incident, even if the zones have physical fire barriers, chances that ope…

True, I implicitly included the MTTR in the "severity", but this is actually a different thing (severity is more about the impact radius). But I don't think it changes my point: knowing what/how Google Cloud designs regions or zones is still an implementation detail, what matters is what MTTR they are targeting and this should be known ahead of time. There are so many "implementation details" that customers are not a…

Scale of impact, scope of impact, and duration of impact are orthogonal. Conflating them makes productive discussion impossible, IMO.

But back to the point, philosophically I agree, but practically I don't. IMO having SLA's and enforceable guarantees that give customers the information they need is much harder than exposing the implementation details.

"Zones within a region may be located in the same building" is much more concise than SLA's using contractual language, and probably conveys more (though potentially less accurate) information once I apply my context.

Also, if we look GCP's SLA's, this outage blew the SLA breach threshold out of the water for many services. Some are pushing 2 9's of downtime from this incident alone.

Finally (in hindsight maybe I should have led with this, but I'm too lazy to restructure this comment), SLA's are a joke. Outages can destroy your business, but all you get from your cloud provider is that they comp you for usually a small fraction of what they charge you. They have no teeth, so if you can't just write off a major outage you have to have a plan to avoid it, which means you need to know the implementation details

Re: Google Cloud Europe service disruption

#153
post #53

Earlier quoted context omitted.

It was probably striking for better working conditions so Google terminated it

A memorable observation from idlewords is that Googlers will organize for better conditions for Googlers but they never organize for better conditions for their users.

Silicon Valley in a nutshell.
Post reply on HN