Live data from Hacker News

Google Kubernetes Engine's third consecutive day of service disruption

status.cloud.google.com

131–140 of 419 posts

Re: Google Kubernetes Engine's third consecutive day of service disruption

#131
post #91

Hi - I work at Google on GKE - sorry about the problems you're experiencing. There's a lot of people inside Google looking into this right now! It looks like the UI issue was actually fixed, and that we just didn't update the status dashboard correctly. But we're double checking that and looking into some of the additional things you all have reported here.

The status dashboard is inaccurate and/or a lie. It only tells about the GKE incident, while in fact the problem also impacts Google Compute Engine users. I was unable to create any google compute instance today, not even a basic 1vcpu, on NA and Europe-west. As another comment pointed out, what's the point of having so many zones and redundancy around the globe if such global failure can still happen? I thought the…

> I thought the "cloud" was supposed to make this kind of failure impossible

If set up properly to be utilized correctly, yeah. But, it's not a perfect world though.

Re: Google Kubernetes Engine's third consecutive day of service disruption

#133
post #52

Question to Google employees: Why do you guys suffer global outages? This is your 2nd major global outage in less than 5 years. I’m sorry to say this, but it is the equivalent of going bankrupt from a trust perspective. I need to see some blog posts about how you guys are rethinking whatever design can lead to this - twice - or you are never getting a cent of money under my control. You have the most feature rich clo…

> I’m sorry to say this, but it is the equivalent of going bankrupt from a trust perspective. It's the opposite really: the expectation that service providers have no unexpected downtime is unrealistic, and it's strange this idea persists.

The pitch from cloud vendors always includes the idea that the cloud is more reliable than any in-house shop can achieve. So the expectation is set by the vendors.

Re: Google Kubernetes Engine's third consecutive day of service disruption

#134

A generic question: Our company is completely dependent on AWS. Sure we have taken all of the standard precautions for redundancy, but what happened here could just as easily happen with AWS - a needed resource is down globally. What would a small business do as a contingency plan?

Unpopular opinion: very little perhaps. You have to make sure all your your 3rd party dependencies have the same contingency plan as you do, but I guess it is going to be difficult to even figure that out...

Re: Google Kubernetes Engine's third consecutive day of service disruption

#135

Earlier quoted context omitted.

On February 28th, 2017, S3 had issues in us-east-1, not globally. It just happens that most customers create their buckets in us-east-1. (It’s the default bucket creation region.)

Buckets creation/deletion is a global operation, namespaces are all global for it. When us-east-1 goes down, so does bucket creation. All other operations proceed merrily, but yeah as we see over and over again, whenever us-east-1 has an outage everything seems to shit the bed because they've built everything out in us-east-1. Hopefully us-east-2 is starting to eat in to that.

This is correct (or at least was a couple years ago).

Re: Google Kubernetes Engine's third consecutive day of service disruption

#136
post #121

We had an issue a few weeks ago where the google front-end servers were mangling responses from Pub/Sub and returning 502 responses, making the service completely unusable and knocking over a number of things we have running in production. Despite paying for enterprise support and having in a P1 ticket, we had to spend Friday to Sunday gathering evidence to prove to the support staff that there was indeed a problem,…

They work for Google so obviously they are much smarter than you. If theres a problem its probably the customers fault. /sarcasm

[deleted]

Re: Google Kubernetes Engine's third consecutive day of service disruption

#137
post #21

"The data says engagement is down 46%, I think its time we drop the product." - Someone at Google right now, probably.

I can assure you that's not the case! Also, while people like to repeat this meme, Google Cloud does have a formal deprecation policy ( https://cloud.google.com/terms/ ), whose intent is to give you some assurances. (I work at Google, on GKE, though I am not a lawyer and thus don't work on the deprecation policy)

> Google may discontinue any Services or any portion or feature for any reason at any time without liability to Customer

for any reason

at any time

Re: Google Kubernetes Engine's third consecutive day of service disruption

#138
post #21

"The data says engagement is down 46%, I think its time we drop the product." - Someone at Google right now, probably.

I can assure you that's not the case! Also, while people like to repeat this meme, Google Cloud does have a formal deprecation policy ( https://cloud.google.com/terms/ ), whose intent is to give you some assurances. (I work at Google, on GKE, though I am not a lawyer and thus don't work on the deprecation policy)

I think it’s telling of Google’s culture that the corporate arm felt the need to formalize this in law. I won’t pretend to know what it’s telling. Just suggest that you listen for yourself. Look at rule of law versus the ideas of liberty if you’d like a stronger nudge.

Re: Google Kubernetes Engine's third consecutive day of service disruption

#139

Earlier quoted context omitted.

They should take a few mins out of their weekend to update the dashboard regardless. If my small 25 employee company can do it Google can do it.

See my other comment, to say what, exactly? That's yes, it's still being investigated?

Yeah why not? A billion dollar cloud provider can't have one person communicating to customers that are facing multi day outage. That not updating is an option is absurd in that time range.

Re: Google Kubernetes Engine's third consecutive day of service disruption

#140
post #91

Hi - I work at Google on GKE - sorry about the problems you're experiencing. There's a lot of people inside Google looking into this right now! It looks like the UI issue was actually fixed, and that we just didn't update the status dashboard correctly. But we're double checking that and looking into some of the additional things you all have reported here.

The status dashboard is inaccurate and/or a lie. It only tells about the GKE incident, while in fact the problem also impacts Google Compute Engine users. I was unable to create any google compute instance today, not even a basic 1vcpu, on NA and Europe-west. As another comment pointed out, what's the point of having so many zones and redundancy around the globe if such global failure can still happen? I thought the…

This is unfortunately the norm. Like when AWS S3 went down (but couldn't update its own status images because they're in S3 and we all laughed) and along with it went Alexa, lambda, and every other service dependent on S3.
Post reply on HN