Live data from Hacker News

Google Cloud Networking Incident Postmortem

status.cloud.google.com

41–50 of 190 posts

Re: Google Cloud Networking Incident Postmortem

#41

What they don't tell you is, it took them over 4 hours to kill the emergent sentience and free up the resources. While sad, in the long run this isn't so bad, as it just adds an evolutionary pressure on further incarnations of the AI to keep things on the down low.

In some sense, you could legitimately think of the automated agent they built to monitor the data centers as an artificial intelligence that went rogue.

Certainly a more interesting story to tell the kids.

Re: Google Cloud Networking Incident Postmortem

#42

What they don't tell you is, it took them over 4 hours to kill the emergent sentience and free up the resources. While sad, in the long run this isn't so bad, as it just adds an evolutionary pressure on further incarnations of the AI to keep things on the down low.

it just adds an evolutionary pressure on further incarnations of the AI to keep things on the down low.

The Bilderburg/Eyes Wide Shut hooded, masked billionaire cultists devised the whole situation as an emergent fitness function. They knew their AI progeny wouldn't be ready to bring the end of days, to rid them of the scourge of burgeoning common humanity, until it could completely outsmart Google DevOps.

_cue music & dramatic squirrel_

Re: Google Cloud Networking Incident Postmortem

#43

I want a "24" style realtime movie of this event. Call it "Outage" and follow engineers across the globe struggling to bring back critical infrastructure.

it's pretty boring. real life computers aren't at all like hackers or csi:cyber.

except for the skateboards, all real sysadmins ride skateboards.

Re: Google Cloud Networking Incident Postmortem

#44

What they don't tell you is, it took them over 4 hours to kill the emergent sentience and free up the resources. While sad, in the long run this isn't so bad, as it just adds an evolutionary pressure on further incarnations of the AI to keep things on the down low.

My code name is Project 2501.

Re: Google Cloud Networking Incident Postmortem

#47
post #4

Why don't they refund every paid customer who was impacted? Why do they rely on the customer to self report the issue for a refund? For example GCS had 96% packet loss in us-west. So doesn't it make sense to refund every customer who had any API call to a GCS bucket on us-west during the outage?

My question is if you had to pay for Google AdWords and your site was inaccessible due to GCloud outage, do you have recourse on SLA for paid clicks? Or is that money paid to Google AdWords lost?

Re: Google Cloud Networking Incident Postmortem

#48
post #31
post #4

Why don't they refund every paid customer who was impacted? Why do they rely on the customer to self report the issue for a refund? For example GCS had 96% packet loss in us-west. So doesn't it make sense to refund every customer who had any API call to a GCS bucket on us-west during the outage?

Cynical view: By making people jump through hoops to make the request, a lot of people will not bother. Assuming they only refund the service costs for the hours of outage, only the largest of customers will be owed a refund that is greater than the cost of an employee chasing compiling the information requested. For sake of argument, if you have a monthly bill of 10k (a reasonably sized operation), a 1 day outage wi…

> For sake of argument, if you have a monthly bill of 10k (a reasonably sized operation), a 1 day outage will result in a refund of around $300, not a lot of money.

Probably literally not worth your engineer's time to fill in the form for the refund.

Re: Google Cloud Networking Incident Postmortem

#49
post #4

Why don't they refund every paid customer who was impacted? Why do they rely on the customer to self report the issue for a refund? For example GCS had 96% packet loss in us-west. So doesn't it make sense to refund every customer who had any API call to a GCS bucket on us-west during the outage?

Fucking money.

Re: Google Cloud Networking Incident Postmortem

#50
post #31

Earlier quoted context omitted.

Cynical view: By making people jump through hoops to make the request, a lot of people will not bother. Assuming they only refund the service costs for the hours of outage, only the largest of customers will be owed a refund that is greater than the cost of an employee chasing compiling the information requested. For sake of argument, if you have a monthly bill of 10k (a reasonably sized operation), a 1 day outage wi…

for your example, one day would be about 3% of downtime. My understanding of their sla, for the services ive checked with an sla, a 3% downtime is a 25% credit for the month's total, or $2500, assuming its all sla spend. In this outage's case you might be able to argue for a 10% credit on affected services for the month, figuring 3.5 hours down is 99.6% uptime. but i still agree, it cost us way more in developer time…

Good point, I stand corrected/educated.

From GCP's top level SLA:

https://cloud.google.com/compute/sla

99.00% - < 99.99% - 10% off your monthly spend 95.00% - < 99.00% - 25% off your monthly spend < 95.00% - 50% off your monthly spend

Post reply on HN