Hi - I work at Google on GKE - sorry about the problems you're experiencing. There's a lot of people inside Google looking into this right now! It looks like the UI issue was actually fixed, and that we just didn't update the status dashboard correctly. But we're double checking that and looking into some of the additional things you all have reported here.
The status dashboard is inaccurate and/or a lie. It only tells about the GKE incident, while in fact the problem also impacts Google Compute Engine users. I was unable to create any google compute instance today, not even a basic 1vcpu, on NA and Europe-west. As another comment pointed out, what's the point of having so many zones and redundancy around the globe if such global failure can still happen? I thought the…
Google Kubernetes Engine's third consecutive day of service disruption
111–120 of 419 posts
Re: Google Kubernetes Engine's third consecutive day of service disruption
#112"The data says engagement is down 46%, I think its time we drop the product." - Someone at Google right now, probably.
I can assure you that's not the case! Also, while people like to repeat this meme, Google Cloud does have a formal deprecation policy ( https://cloud.google.com/terms/ ), whose intent is to give you some assurances. (I work at Google, on GKE, though I am not a lawyer and thus don't work on the deprecation policy)
Re: Google Kubernetes Engine's third consecutive day of service disruption
#113Earlier quoted context omitted.
The status dashboard is inaccurate and/or a lie. It only tells about the GKE incident, while in fact the problem also impacts Google Compute Engine users. I was unable to create any google compute instance today, not even a basic 1vcpu, on NA and Europe-west. As another comment pointed out, what's the point of having so many zones and redundancy around the globe if such global failure can still happen? I thought the…
> I was unable to create any google compute instance today, not even a basic 1vcpu, on NA and Europe-west. I've been creating GCP instances in us-central1-a and us-central1-c today without issue. Which zone were you using in NA? I have been noticing unusual restarts, but I haven't been able to pin down the cause yet (may be my software and not GCP itself).
Re: Google Kubernetes Engine's third consecutive day of service disruption
#114Earlier quoted context omitted.
First, determine your tolerance. How much does downtime per min cost you? How long can you be down? A lot of the time it may be cheaper to apologize to your customers than build a truly reliable system. Then start looking at points of failure and sort them based on severity and probability. Is your own software deployment going to generate more downtime per year than a regional aws outage? There are formal academic w…
If all your competitors are down too, the equation changes, no?
Any trade comes to me if it's urgent, and I appear more professional as I've got a functioning system.
I might be an chancer running my entire system off an shoestring but being up when everyone else has taken a dive looks good.
Re: Google Kubernetes Engine's third consecutive day of service disruption
#115Earlier quoted context omitted.
It's weekend, why wouldn't you take it off? It's just silly software.
A lot of people's businesses, reputations and livelihoods depend on "silly software," not to mention that they are paying customers themselves.
Re: Google Kubernetes Engine's third consecutive day of service disruption
#116Earlier quoted context omitted.
I'm not sure where you get this idea, but AWS has definitely had global outages. 1.5 years ago there were massive global issues with S3, causing even their own status dashboard to malfunction. edit: I stand corrected. Apparently the S3 outage wasn't global, though its effects were. Meanwhile, this outage has only really been noticeable to ops teams, since it doesn't affect existing nodes or anything outside GKE. It's…
On February 28th, 2017, S3 had issues in us-east-1, not globally. It just happens that most customers create their buckets in us-east-1. (It’s the default bucket creation region.)
Re: Google Kubernetes Engine's third consecutive day of service disruption
#117Earlier quoted context omitted.
Infrastructure as code. Terraform using AMIs plus chef recipes that work in the cloud and bare metal. Dont use AWS specific services. This would allow you to spin over to another cloud provider , vsphere or bare metal with minimal work
I think you are downplaying minimal
Even when working in small companies with small infrastructure, I've kept recreation of infrastructure as one of my high priorities (one reason it really bugged me in one job to have to depend on Oracle Databases that I couldn't automate to the same degree.)
In my mind, it's not different from the importance of having, and testing restoration of, backups. If your infrastructure gets compromised somehow, or you find yourself up the creek with your provider, you've got to be able to rebuild everything from scratch.
Re: Google Kubernetes Engine's third consecutive day of service disruption
#118Re: Google Kubernetes Engine's third consecutive day of service disruption
#119Earlier quoted context omitted.
Understandable but in my experience the incident manager assigned is still supposed to keep track of progress during weekends when you have a major incident. And this is likely a major incident with significant customer impact. The way google is handling all this gives a pretty poor impression. Seems like this kubernetes is just a PoC.
Incident manager isn't public comms person. The person updating that status dashboard may or may not be an engineer, the IM certainly is.
Re: Google Kubernetes Engine's third consecutive day of service disruption
#120Earlier quoted context omitted.
Understandable but in my experience the incident manager assigned is still supposed to keep track of progress during weekends when you have a major incident. And this is likely a major incident with significant customer impact. The way google is handling all this gives a pretty poor impression. Seems like this kubernetes is just a PoC.
Incident manager isn't public comms person. The person updating that status dashboard may or may not be an engineer, the IM certainly is.