Live data from Hacker News

Google Kubernetes Engine's third consecutive day of service disruption

status.cloud.google.com

401–410 of 419 posts

Re: Google Kubernetes Engine's third consecutive day of service disruption

#401

Earlier quoted context omitted.

Oh, courts routinely give binding weight to words like Google's deprecation policy uses, and any large megacorp who is sufficiently badly impacted by a legalese violation (though SLA issues and deprecation issues are two separate things) wouldn't be scared away from a lawsuit by Google being big. I can imagine EU regulatory action or a class action lawsuit as other possible mechanisms. But as I say in another comment…

> The reality of what Google has done and will do with GCP, though, is pretty good. No. It's just words. Actions speak louder than words. Googles' actions in the last couple of days spoke pretty loud. No amount of words will change that. Are you working for Google PR or something?

I haven't worked for Google since 2015, and I never worked for their PR department. I was just a rank-and-file engineer (and a rank-and-file tech lead for one small team near the end of my time there). If I worked for Google PR, my comments throughout this thread would have far less criticism of the company's messaging and branding than they do. :)

I'm still a fan of GCP as a suite of products and services, as much as I recognize many of Google's organizational failings and disagree with plenty of their product decisions in other areas of Google.

Google (including GCP) has been bad at external communication as long as I've paid attention, and that includes external communications around incidents. What actions are you referring to, beyond poor and confusing communication (i.e. words) around what is or isn't broken or fixed at what points during the incident? That's most of the problem I'm aware of from this incident.

With that said, part of the reason people notice GCP's outages more than AWS's is that GCP publicly notes their outages way more than AWS does. In other words, among the outages that either cloud has, Google much more often creates an incident on their public status page and Amazon much more often fails to.

My "reality of [...] GCP" comment was about the bigger picture of the cloud platform offering, not any one specific incident.

Re: Google Kubernetes Engine's third consecutive day of service disruption

#402
post #48

A generic question: Our company is completely dependent on AWS. Sure we have taken all of the standard precautions for redundancy, but what happened here could just as easily happen with AWS - a needed resource is down globally. What would a small business do as a contingency plan?

We invested early in being multi-region on GCP as well as multi-cloud with AWS as a fully redundant option if it ever became necessary to fail over to them.

Paying off big-time now :-)

Re: Google Kubernetes Engine's third consecutive day of service disruption

#403
post #402
post #48

Earlier quoted context omitted.

We invested early in being multi-region on GCP as well as multi-cloud with AWS as a fully redundant option if it ever became necessary to fail over to them.

Paying off big-time now :-)

It sure did. It took us about a minute to fully fail over to AWS, and we are running there now for the rest of the day. We will promote GCP to be the primary cloud again tonight.

Without our multi-cloud set up we would have been down for over an hour. In our business this is not an option.

Re: Google Kubernetes Engine's third consecutive day of service disruption

#404

Earlier quoted context omitted.

Your response is from an IT consumer's viewpoint. "Cloud" is not a thing one buys and one's reputation has nothing to do with the reliability of the services consumed, but the reliability of the services provided. To put it more succinctly, "you own your availability". In the end, "cloud" is a commodity and all cloud providers are trying to get vendor lock-in. My goal as a manager is not to couple my business revenue…

So you’re not using any third party vendor for anything and you’re doing everything in house? Cloud is only an interchangeable commodity if you’re treating it like an overpriced colo and not using it to save costs on staff, maintenance, and helping deliver product faster.

That's why it's important to use k8s (and various abstraction layers, standards APIs, etc.), which helps with portability, and thus you can risk load balance between cloud vendors, and you can even throw in on-prem into the mix.

Re: Google Kubernetes Engine's third consecutive day of service disruption

#405
post #404

Earlier quoted context omitted.

So you’re not using any third party vendor for anything and you’re doing everything in house? Cloud is only an interchangeable commodity if you’re treating it like an overpriced colo and not using it to save costs on staff, maintenance, and helping deliver product faster.

That's why it's important to use k8s (and various abstraction layers, standards APIs, etc.), which helps with portability, and thus you can risk load balance between cloud vendors, and you can even throw in on-prem into the mix.

And why wouldn't you just do a colo then? From my experience, cloud infrastructure is always more expensive than the equivalent on prem/colo infrastructure unless you depend on hosted solutions and you're willing and able to operate with fewer infrastructure people and it doesn't help you move faster.

Re: Google Kubernetes Engine's third consecutive day of service disruption

#406
post #404

Earlier quoted context omitted.

That's why it's important to use k8s (and various abstraction layers, standards APIs, etc.), which helps with portability, and thus you can risk load balance between cloud vendors, and you can even throw in on-prem into the mix.

And why wouldn't you just do a colo then? From my experience, cloud infrastructure is always more expensive than the equivalent on prem/colo infrastructure unless you depend on hosted solutions and you're willing and able to operate with fewer infrastructure people and it doesn't help you move faster.

The last time I did the math a reasonable highly available setup was $1-3m CapEx per data center, I’d want no less than three. That’s 30kw worth of gear per for a total of 90kw at ~$180/kw MRC if I’m lucky plus transit fees, so $20k a month.

Doable, but it’s a hell of a lot of hassle and that CapEx is huge for a startup.

I’d go bare metal in a second for any kind of cost conscious business that needed scale and had established revenue.

Re: Google Kubernetes Engine's third consecutive day of service disruption

#407

Earlier quoted context omitted.

And why wouldn't you just do a colo then? From my experience, cloud infrastructure is always more expensive than the equivalent on prem/colo infrastructure unless you depend on hosted solutions and you're willing and able to operate with fewer infrastructure people and it doesn't help you move faster.

The last time I did the math a reasonable highly available setup was $1-3m CapEx per data center, I’d want no less than three. That’s 30kw worth of gear per for a total of 90kw at ~$180/kw MRC if I’m lucky plus transit fees, so $20k a month. Doable, but it’s a hell of a lot of hassle and that CapEx is huge for a startup. I’d go bare metal in a second for any kind of cost conscious business that needed scale and had e…

In the case of a startup, the question is just the opposite. It’s about moving fast and being nimble more than it is about worrying about a distant future where lock-in could possibly be an issue. I would be all in on leveraging as many of the services that the cloud provider offers and where they could take care of the “undifferentiated heavy lifting”.

Re: Google Kubernetes Engine's third consecutive day of service disruption

#408

Earlier quoted context omitted.

10x engineers are real. Start-up geniuses are real. The large majority of people have their heads up their asses. Wake up and smell the coffee

You can't find and hire geniuses for every component. Systems should scale with the average engineer in mind (so should code). We have all smelt the coffee and it smells even better when your team is efficient and well rested.

I don't disagree. Building around 10x'ers in this way creates god complexes and unhappy "senior" engineers (since they don't do any of the cool work). Having one very early still juices your productivity

Re: Google Kubernetes Engine's third consecutive day of service disruption

#409
post #250

Earlier quoted context omitted.

- Native integration with G-Suite as an identity provider. Unified permissions modeling from the IDP, to work apps like email/Drive, to cloud resources, all the way into Kubernetes IAM. - Security posture. Project Zero is class leading, and there's absolutely a "fear-based" component there, with the open question of when Project Zero discovers a new exploit, who will they share it with before going public? The upcomi…

> If you're on Kubernetes and its not on GKE, then you've got legacy reasons for being where you're at. Or you have legitimate reasons for running on your own hardware, e.g. compliance or locality (I work at SAP's internal cloud and we have way more regions than the hyperscalers because our customers want to have their data stay in their own country).

That's totally fair; that comment was just in reference to comparing other cloud providers (AWS and Azure primarily)

Re: Google Kubernetes Engine's third consecutive day of service disruption

#410
post #365

Earlier quoted context omitted.

Quit using vendor lock in resources then asking why its hard to leave that vendor.

Oh, I see. You intend to use purely EC2 or GCE instances and just run everything one would run in a colo on them. All right. That's a pretty manpower intensive way of operating. I think the fact that you get cloud agnostic this way is probably not worth it.

The requirement was to be cloud agnostic. This is one of the few ways possible.
Post reply on HN