Live data from Hacker News

Google Kubernetes Engine's third consecutive day of service disruption

status.cloud.google.com

91–100 of 419 posts

Re: Google Kubernetes Engine's third consecutive day of service disruption

#91

Hi - I work at Google on GKE - sorry about the problems you're experiencing. There's a lot of people inside Google looking into this right now! It looks like the UI issue was actually fixed, and that we just didn't update the status dashboard correctly. But we're double checking that and looking into some of the additional things you all have reported here.

The status dashboard is inaccurate and/or a lie. It only tells about the GKE incident, while in fact the problem also impacts Google Compute Engine users. I was unable to create any google compute instance today, not even a basic 1vcpu, on NA and Europe-west.

As another comment pointed out, what's the point of having so many zones and redundancy around the globe if such global failure can still happen? I thought the "cloud" was supposed to make this kind of failure impossible

Re: Google Kubernetes Engine's third consecutive day of service disruption

#92
I honestly don't mind if providers have outages - we can't expect 100.00% accuracy, I know the systems I manage certainly don't achieve that.

One thing I do care about though, is root cause analysis. I love reading a good RCA, it restores my faith in the company and makes me trust them more.

(I'm not affect by the GKE outage so opinions may differ right now!)

Re: Google Kubernetes Engine's third consecutive day of service disruption

#93

Earlier quoted context omitted.

The people tasked with fixing this aren't the ones providing the updates.

Understandable but in my experience the incident manager assigned is still supposed to keep track of progress during weekends when you have a major incident. And this is likely a major incident with significant customer impact. The way google is handling all this gives a pretty poor impression. Seems like this kubernetes is just a PoC.

Incident manager isn't public comms person.

The person updating that status dashboard may or may not be an engineer, the IM certainly is.

Re: Google Kubernetes Engine's third consecutive day of service disruption

#94

A generic question: Our company is completely dependent on AWS. Sure we have taken all of the standard precautions for redundancy, but what happened here could just as easily happen with AWS - a needed resource is down globally. What would a small business do as a contingency plan?

You need to have a secondary location for backups, not AWS. Stores copie of customer data, orders, accounts, balances and anything that is critical to the business.

If AWS ever screws up, you will be able to continue running the business even if it might take weeks to start over.

For live redundancy, you should have a secondary datacenter on another provider, but realistically it's hard to do and most business never achieve that. Instead, just stick with AWS and if there is a problem the strategy is to sip coffee while waiting for them to resolve it. Much better this way than you having to fix it yourself.

Re: Google Kubernetes Engine's third consecutive day of service disruption

#95

I have a question. At what point does k8s make sense? I have a feeling that a microservice architecture is overkill for 99% of businesses. You can serve a lot of customers on a single node with the hardware available today. Often times, sharding on customers is rather trivial as well. Monolith for the win! Opinions?

There's a huge range between monolith and microservice approach, and even a monolith will have dependent services. A simple web stack these days might include nginx, a database, a caching layer, some sort of task broker and then the 'monolith' web app itself. All of that can be sanely managed in k8s.

Re: Google Kubernetes Engine's third consecutive day of service disruption

#96
post #59

Earlier quoted context omitted.

2 outage in 5 years sounds pretty low, to be honest. Disclaimer: google employee in ads, who worked on many many fires throughout the years, but talking from my personal perspective and not from my employer. I am sure we are striving to have 0, but realistically, i have seen many that says things happen. Learn, and improve.

The issue people have with it is that it's global, not regional, indicating that there are dependencies in the entire architecture that people does not expect to be there.

Yes, hello, canaries anyone

Re: Google Kubernetes Engine's third consecutive day of service disruption

#97
post #86
post #45

Earlier quoted context omitted.

Yeah, AWS has never had a global outage, Google has had 2 now.

I'm not sure where you get this idea, but AWS has definitely had global outages. 1.5 years ago there were massive global issues with S3, causing even their own status dashboard to malfunction. edit: I stand corrected. Apparently the S3 outage wasn't global, though its effects were. Meanwhile, this outage has only really been noticeable to ops teams, since it doesn't affect existing nodes or anything outside GKE. It's…

On February 28th, 2017, S3 had issues in us-east-1, not globally. It just happens that most customers create their buckets in us-east-1. (It’s the default bucket creation region.)

Re: Google Kubernetes Engine's third consecutive day of service disruption

#98

Earlier quoted context omitted.

Fair point but still seems odd that the people providing updates took the weekend off during a large scale customer impacting issue. I'm sure all the people spending the weekend trying to mitigate the impact of this on their infrastructure would love to have timely updates.

This is a problem in the web ui creating additional node pools, there is a very simple workaround to use gcloud. What major impact are you referring to? It's likely the fix is checked in and will start roll out on Monday. Disclaimer: I work on Google Cloud and while I believe we could use more words here, this doesn't seem like a huge problem. It's embarrassing the the issue with the ui was shipped, and I'm sure this…

It's not just the UI though. Resources have been exhausted, or that's the error I'm getting, in a lot of regions for both my work and personal accounts (GKE and GCE). Someone on the GCP slack also said they were getting similar issues from Dataflow so it seems to be widespread across products.

Re: Google Kubernetes Engine's third consecutive day of service disruption

#99
post #91

Hi - I work at Google on GKE - sorry about the problems you're experiencing. There's a lot of people inside Google looking into this right now! It looks like the UI issue was actually fixed, and that we just didn't update the status dashboard correctly. But we're double checking that and looking into some of the additional things you all have reported here.

The status dashboard is inaccurate and/or a lie. It only tells about the GKE incident, while in fact the problem also impacts Google Compute Engine users. I was unable to create any google compute instance today, not even a basic 1vcpu, on NA and Europe-west. As another comment pointed out, what's the point of having so many zones and redundancy around the globe if such global failure can still happen? I thought the…

> I was unable to create any google compute instance today, not even a basic 1vcpu, on NA and Europe-west.

I've been creating GCP instances in us-central1-a and us-central1-c today without issue. Which zone were you using in NA?

I have been noticing unusual restarts, but I haven't been able to pin down the cause yet (may be my software and not GCP itself).

Re: Google Kubernetes Engine's third consecutive day of service disruption

#100
post #44

Earlier quoted context omitted.

First, determine your tolerance. How much does downtime per min cost you? How long can you be down? A lot of the time it may be cheaper to apologize to your customers than build a truly reliable system. Then start looking at points of failure and sort them based on severity and probability. Is your own software deployment going to generate more downtime per year than a regional aws outage? There are formal academic w…

People outsource to cloud providers because building / hiring / maintaining a team of decent engineers that provide a baseline industry bar of SLAs, SLOs is much more expensive than the eye watering costs of most cloud providers at even a IaaS level. Opex is tough. Most companies I’ve been at don’t offer multi region support for their services because it’s too expensive for the service provided even in so-called “pri…

The baseline is that it takes 12 dedicated people across the world to run a 24/7 support operation.

Considering that even tech companies hardly manage to have a pair of DevOps or Sysadmin, running one own infrastructure is completely out of question.

Post reply on HN