Live data from Hacker News

Google Kubernetes Engine's third consecutive day of service disruption

status.cloud.google.com

391–400 of 419 posts

Re: Google Kubernetes Engine's third consecutive day of service disruption

#391
As of this morning, I am still unable to reliably start my docker+machine autoscaling instances. In all cases the error is "Error: The zone does not have enough resources available to fulfill the request"

An instance in us-central1-a has refused to start since last Thursday or Friday.

I created a new instance in us-west2-c, which worked briefly but began to fail midday Friday, and kept failing through the weekend.

On Saturday I created yet another clone in northamerica-northeast1-b. That worked Saturday and Sunday, but this morning, it is failing to start. Fortunately my us-west2-c instance has begun to work again, but I'm having doubts about continuing to use GCE as we scale up.

And yet, the status page says all services are available.

Is the typical of others' experiences?

Re: Google Kubernetes Engine's third consecutive day of service disruption

#392

Hi - I work at Google on GKE - sorry about the problems you're experiencing. There's a lot of people inside Google looking into this right now! It looks like the UI issue was actually fixed, and that we just didn't update the status dashboard correctly. But we're double checking that and looking into some of the additional things you all have reported here.

As of this morning, I am still unable to reliably start my docker+machine autoscaling instances. In all cases the error is "Error: The zone does not have enough resources available to fulfill the request" An instance in us-central1-a has refused to start since last Thursday or Friday.

I created a new instance in us-west2-c, which worked briefly but began to fail midday Friday, and kept failing through the weekend.

On Saturday I created yet another clone in northamerica-northeast1-b. That worked Saturday and Sunday, but this morning, it is failing to start. Fortunately my us-west2-c instance has begun to work again, but I'm having doubts about continuing to use GCE as we scale up.

And yet, the status page says all services are available.

Re: Google Kubernetes Engine's third consecutive day of service disruption

#394

I am currently evaluating GCP for two separate projects. I want to see if I understand this correctly: 1) For three whole days, it was questionable whether or not a user would be able to launch a node pool (according to the official blog statement). It was also questionable whether a user would be able to launch a simple compute instance (according to statements here on HN). 2) This issue was global in scope, affecti…

You're doing me a scare. I'm in the evaluation phase with them. Maybe I'm missing something here, but this is not at all what the linked post says. "We are investigating an issue with Google Kubernetes Engine node pool creation through Cloud Console UI." So, it's a UI console issue, it appears you can still manage "Affected customers can use gcloud command [1] in order to create new Node Pools. [1]" Similarly, it act…

As of this morning, I am still unable to reliably start my docker+machine autoscaling instances. In all cases the error is "Error: The zone does not have enough resources available to fulfill the request" An instance in us-central1-a has refused to start since last Thursday or Friday.

I created a new instance in us-west2-c, which worked briefly but began to fail midday Friday, and kept failing through the weekend.

On Saturday I created yet another clone in northamerica-northeast1-b. That worked Saturday and Sunday, but this morning, it is failing to start. Fortunately my us-west2-c instance has begun to work again, but I'm having doubts about continuing to use GCE as we scale up.

And yet, the status page says all services are available.

Re: Google Kubernetes Engine's third consecutive day of service disruption

#395
post #390

Earlier quoted context omitted.

Other than kubernetes and object storage, no.

It could work then. We had the same setup. Also, we store everything on S3 but there is a lambda function that pushes new objects to Google Cloud storage too. We figured, storage is cheap so we can duplicate things and it will make moving to GCP a lot easier. This was a while back though. Now we depend on a lot more AWS stuff.

At one point, I was maintaining clusters in AWS, Azure, and GCP. It was a more work than I anticipated, for very little benefit.

Cluster administration and identity management are unique to each provider and fairly challenging to get right.

Re: Google Kubernetes Engine's third consecutive day of service disruption

#396
post #365

Earlier quoted context omitted.

Assume I’m not fully informed here. What does “written properly” mean? Sure I can move over route53 to cloud DNS easily but Firehose to PubSub? Lambda to Cloud Functions? DynamoDB to BigTable and moving the data? The syntax for provisioning these doesn’t work that well for some find and replace to work. Are you using a templater to generate cloud-specific HCL from a template or something? Sounds like a pretty big pro…

Quit using vendor lock in resources then asking why its hard to leave that vendor.

Oh, I see. You intend to use purely EC2 or GCE instances and just run everything one would run in a colo on them. All right.

That's a pretty manpower intensive way of operating. I think the fact that you get cloud agnostic this way is probably not worth it.

Re: Google Kubernetes Engine's third consecutive day of service disruption

#397

Earlier quoted context omitted.

Your response is from a geek’s viewpoint. No insult, intended, I’m first and foremost a 30 year computer geek myself - started programming in 65C02 assembly in 6th grade and still mostly hands on. But, whether it is right or not, as an architect/manager, etc, you have to think about what’s not just best technically. You also have to manage your reputational risks if things go south and less selfishly, how quickly can…

Your response is from an IT consumer's viewpoint. "Cloud" is not a thing one buys and one's reputation has nothing to do with the reliability of the services consumed, but the reliability of the services provided. To put it more succinctly, "you own your availability". In the end, "cloud" is a commodity and all cloud providers are trying to get vendor lock-in. My goal as a manager is not to couple my business revenue…

So you’re not using any third party vendor for anything and you’re doing everything in house?

Cloud is only an interchangeable commodity if you’re treating it like an overpriced colo and not using it to save costs on staff, maintenance, and helping deliver product faster.

Re: Google Kubernetes Engine's third consecutive day of service disruption

#398

Earlier quoted context omitted.

So basically they screw you over on either medium or hard, and according to you the real problem is them not telling us clearly whether or not we receive the medium or hard screw over option. By the way that GCP is so full of loopholes where Google can get out of its obligations its laughable. So it's not even that clear cut that the GCP is really a better alternative. And even when it turns out to be legally sound,…

Oh, courts routinely give binding weight to words like Google's deprecation policy uses, and any large megacorp who is sufficiently badly impacted by a legalese violation (though SLA issues and deprecation issues are two separate things) wouldn't be scared away from a lawsuit by Google being big. I can imagine EU regulatory action or a class action lawsuit as other possible mechanisms. But as I say in another comment…

> The reality of what Google has done and will do with GCP, though, is pretty good.

No. It's just words. Actions speak louder than words. Googles' actions in the last couple of days spoke pretty loud. No amount of words will change that.

Are you working for Google PR or something?

Re: Google Kubernetes Engine's third consecutive day of service disruption

#399

Earlier quoted context omitted.

How could I be incorrect when that's exactly what I said? You gotta pay for a support contract to have any meaningful support.

You also said that all you would get was an "automated ack". This seems to not be the case if aws provides an on-call support engineer.

I think the point is that that only happens if you have a contract. With GCP you can also get an oncall support engineer if you're large enough.

Re: Google Kubernetes Engine's third consecutive day of service disruption

#400

Earlier quoted context omitted.

When everything works, GCP is the best. Stable, fast, simple, reliable. When things stop working, GCP is the worst. Slow communications and they require way too much work before escalating issues or attempting to find a solution. They already have the tools and access so most issues should take minutes for them to gather diagnostics, but instead they keep sending tickets back for "more info", inevitably followed by a…

> instead they keep sending tickets back for "more info" Isn't that the case with basically every support request, no matter the company or severity? The first couple of emails from 1st & even 2nd level support are mostly about answering the same questions about the environment over and over again. We've had this ping-pong situation with production outages (which we eventually analysed and worked around by ourselves)…

Smaller companies or personalized support structures (like named engineers) is very different. You can build up a relationship and usually bypass many questions to get to main issue, and even get it resolved before you can even open a case at larger organizations.

GCP does have role-based support models with a flat-rate plan, which is really great, but the overall quality of the responses leaves much to be desired.

Post reply on HN