Live data from Hacker News

Google Cloud Networking Incident Postmortem

status.cloud.google.com

1–10 of 190 posts

Re: Google Cloud Networking Incident Postmortem

#2
Having only ever seen one major outage event in person (at a financial institution that hadn't yet come up with an incident response plan; cue three days of madness), I would love to be a fly on the wall at Google or other well-established engineering orgs when something like this goes down.

I'd love to see the red binders come down off the shelf, people organize into incident response groups, and watch as a root cause is accurately determined and a fix out in place.

I know it's probably more chaos than art, but I think there would be a lot to learn by seeing it executed well.

Re: Google Cloud Networking Incident Postmortem

#3
What they don't tell you is, it took them over 4 hours to kill the emergent sentience and free up the resources. While sad, in the long run this isn't so bad, as it just adds an evolutionary pressure on further incarnations of the AI to keep things on the down low.

Re: Google Cloud Networking Incident Postmortem

#4
Why don't they refund every paid customer who was impacted? Why do they rely on the customer to self report the issue for a refund?

For example GCS had 96% packet loss in us-west. So doesn't it make sense to refund every customer who had any API call to a GCS bucket on us-west during the outage?

Re: Google Cloud Networking Incident Postmortem

#5
post #4

Why don't they refund every paid customer who was impacted? Why do they rely on the customer to self report the issue for a refund? For example GCS had 96% packet loss in us-west. So doesn't it make sense to refund every customer who had any API call to a GCS bucket on us-west during the outage?

Probably because it seems to be in the SLA that the customer must notify Google? https://cloud.google.com/storage/sla

> "[Customer Must Request Financial Credit] In order to receive any of the Financial Credits described above, Customer must notify Google technical support within thirty days from the time Customer becomes eligible to receive a Financial Credit. Failure to comply with this requirement will forfeit Customer’s right to receive a Financial Credit."

Re: Google Cloud Networking Incident Postmortem

#6
post #4

Why don't they refund every paid customer who was impacted? Why do they rely on the customer to self report the issue for a refund? For example GCS had 96% packet loss in us-west. So doesn't it make sense to refund every customer who had any API call to a GCS bucket on us-west during the outage?

I agree, but it’s pretty standard SLA verbiage (from the telco/bandwith provider days) to require the customer to request/register the SLA violation to benefit.

Re: Google Cloud Networking Incident Postmortem

#7
> Google Cloud instances in us-west1, and all European regions and Asian regions, did not experience regional network congestion.

Does not appear to be true. Tests I was running on cloud functions in europe-west2 saw impact to europe-west2 GCS buckets.

https://medium.com/lightstephq/googles-june-2nd-outage-their...

Re: Google Cloud Networking Incident Postmortem

#8
post #4

Why don't they refund every paid customer who was impacted? Why do they rely on the customer to self report the issue for a refund? For example GCS had 96% packet loss in us-west. So doesn't it make sense to refund every customer who had any API call to a GCS bucket on us-west during the outage?

Microsoft refunded after their latest outage in South Central. Google might announce a refund later, though I did read on here that some of their outage was not covered by their SLA.

Re: Google Cloud Networking Incident Postmortem

#9
post #4

Why don't they refund every paid customer who was impacted? Why do they rely on the customer to self report the issue for a refund? For example GCS had 96% packet loss in us-west. So doesn't it make sense to refund every customer who had any API call to a GCS bucket on us-west during the outage?

They write the need for the customer to request it into the SLA, probably on the theory that a lot of customers won’t, which saves them money.

Re: Google Cloud Networking Incident Postmortem

#10

What they don't tell you is, it took them over 4 hours to kill the emergent sentience and free up the resources. While sad, in the long run this isn't so bad, as it just adds an evolutionary pressure on further incarnations of the AI to keep things on the down low.

Birth is always traumatic.
Post reply on HN