Live data from Hacker News

Google Kubernetes Engine's third consecutive day of service disruption

status.cloud.google.com

181–190 of 419 posts

Re: Google Kubernetes Engine's third consecutive day of service disruption

#181

A generic question: Our company is completely dependent on AWS. Sure we have taken all of the standard precautions for redundancy, but what happened here could just as easily happen with AWS - a needed resource is down globally. What would a small business do as a contingency plan?

Multi-Cloud-Service-as-a-Service application redundancy wrappers.

I wish I was only being tongue-in-cheek.

Re: Google Kubernetes Engine's third consecutive day of service disruption

#182
post #159
post #155

Earlier quoted context omitted.

It's ok I guess but still lets them turn it off if, in their judgement, its an economic burden i.e. costing them money. If they ever do deprecate something people have built on though they're gonna get absolutely crucified. That's probably better protection than any terms of service.

> It's ok I guess but still lets them turn it off if, in their judgement, its an economic burden i.e. costing them money. If a service Google runs is losing money, what reason would they have to not shut it down?

With this terms of service, none. Which is why people don't trust them.

If I pay you for a service that would take time to migrate off of, and you are making money off me now, I am going to be ripshit if you decide to just turn it off because it's suddenly not making money for you in the short term. Google's done this a lot, and the fact that don't provide concrete time lines in their contract gives even less reason to trust them

Re: Google Kubernetes Engine's third consecutive day of service disruption

#183
post #52

Question to Google employees: Why do you guys suffer global outages? This is your 2nd major global outage in less than 5 years. I’m sorry to say this, but it is the equivalent of going bankrupt from a trust perspective. I need to see some blog posts about how you guys are rethinking whatever design can lead to this - twice - or you are never getting a cent of money under my control. You have the most feature rich clo…

Most feature rich cloud? I think that title belongs to AWS.

Yeah. Perhaps feature rich was an overstatement. I meant that when GCP does do a product it works like I’d expect it to work and has the features I need. Not always the case with a AWS, particularly around ELBs and VPCs.

Re: Google Kubernetes Engine's third consecutive day of service disruption

#184
post #52

Question to Google employees: Why do you guys suffer global outages? This is your 2nd major global outage in less than 5 years. I’m sorry to say this, but it is the equivalent of going bankrupt from a trust perspective. I need to see some blog posts about how you guys are rethinking whatever design can lead to this - twice - or you are never getting a cent of money under my control. You have the most feature rich clo…

I’d be curious to know what alternatives are you considering at this point?

Azure and AWS.

Re: Google Kubernetes Engine's third consecutive day of service disruption

#185
post #59

Earlier quoted context omitted.

2 outage in 5 years sounds pretty low, to be honest. Disclaimer: google employee in ads, who worked on many many fires throughout the years, but talking from my personal perspective and not from my employer. I am sure we are striving to have 0, but realistically, i have seen many that says things happen. Learn, and improve.

The issue people have with it is that it's global, not regional, indicating that there are dependencies in the entire architecture that people does not expect to be there.

There are many other possible causes for global outages, that specific one is not high on my list of likely culprits.

Re: Google Kubernetes Engine's third consecutive day of service disruption

#186

Earlier quoted context omitted.

The baseline is that it takes 12 dedicated people across the world to run a 24/7 support operation. Considering that even tech companies hardly manage to have a pair of DevOps or Sysadmin, running one own infrastructure is completely out of question.

What's the math/logic to get to 12 people?

Two people per 6-hour timezone chunk, with one overlapping person splitting each chunk?

2 at UTC 0-8

1 at UTC 4-12

2 at UTC 8-16

1 at UTC 12-20

2 at UTC 16-24

1 at UTC 20-04

(Repeat)

Re: Google Kubernetes Engine's third consecutive day of service disruption

#187

Earlier quoted context omitted.

But what do these things mean? > commercially reasonable > substantial economic or material technical burden Is one engineer working on an old service to keep it alive commercially reasonable or a substantial burden? I don't know. Do you? In practice this policy lets them shut off anything they want any time they want. Again it's their playground they can do what they want unless they signed a contract saying they'd…

I think you're ascribing an unreasonable amount of bad faith here, and, to rephrase what I had here before, you're approaching this from an engineering perspective, not a legal one. And that's not how those things work. To be clear, that policy is a contract. And those things would be decided by a jury. And if my understanding is correct, the reasonable person standard applies. So you can answer this yourself, do you…

>If not, why mention it?

Because it makes more people feel comfortable enough to use your services and pay you, without actually binding you towards any sort of behavior that would cost you money. There's a direct financial incentive here to use legalese to give the semblance of reliability without having to deliver on it

Re: Google Kubernetes Engine's third consecutive day of service disruption

#188
post #159
post #155

Earlier quoted context omitted.

It's ok I guess but still lets them turn it off if, in their judgement, its an economic burden i.e. costing them money. If they ever do deprecate something people have built on though they're gonna get absolutely crucified. That's probably better protection than any terms of service.

> It's ok I guess but still lets them turn it off if, in their judgement, its an economic burden i.e. costing them money. If a service Google runs is losing money, what reason would they have to not shut it down?

Customers won't pay money in the first place to use a service if it may vanish out from under them? I expect a cloud service provider not to offer a service unless they think it is going to be profitable, and I expect them to continue to offer it even if it turns out not to be profitable, because otherwise I will take my business to a cloud service provider that will give me that guarantee.

Re: Google Kubernetes Engine's third consecutive day of service disruption

#189

I am currently evaluating GCP for two separate projects. I want to see if I understand this correctly: 1) For three whole days, it was questionable whether or not a user would be able to launch a node pool (according to the official blog statement). It was also questionable whether a user would be able to launch a simple compute instance (according to statements here on HN). 2) This issue was global in scope, affecti…

You're doing me a scare. I'm in the evaluation phase with them. Maybe I'm missing something here, but this is not at all what the linked post says.

"We are investigating an issue with Google Kubernetes Engine node pool creation through Cloud Console UI."

So, it's a UI console issue, it appears you can still manage

"Affected customers can use gcloud command [1] in order to create new Node Pools. [1]"

Similarly, it actually was resolved in Friday, but they forgot to mark it as so.

"The issue with Google Kubernetes Engine Node Pool creation through the Cloud Console UI had been resolved as of Friday, 2018-11-09 14:30 US/Pacific."

Re: Google Kubernetes Engine's third consecutive day of service disruption

#190
post #66

Earlier quoted context omitted.

(disclaimer: I work for another cloud provider) I agree, in general, outages are almost inevitable, but global outages shouldn't occur. It suggests at least a couple of things: 1) Bad software deployments, without proper validation. A message elsewhere in this post on HN suggest that problems have been occurring for at least 5 days, which makes me think this is the most likely situation. If this is the case, presumab…

That shouldn't, but they do. S3 goes down [1]. The AWS global console goes down, right after Prime Day outages [2]. Lots of Google Cloud services go down [3, current thread]. Tens of Azure services go down hard [4]. Are software development and release processes improving to mitigate these outages? We don't know. You have to trust the marketing. Will regions ever be fully isolated? We don't know. Will AWS IAM and con…

Look how frequent and detailed Amazon's update logs are in that first Register article. Multiple updates throughout the day going into some detail.
Post reply on HN