Live data from Hacker News

Google Kubernetes Engine's third consecutive day of service disruption

status.cloud.google.com

351–360 of 419 posts

Re: Google Kubernetes Engine's third consecutive day of service disruption

#351
post #5
post #2

Status page is inaccurate as issues doesn't only affect the web UI, the same operations are not functioning via the CLI.

Its kinda strange that HN seems to be the most effective way to give feedback to Google Cloud :/

I also find it weird that on HN where normally people are very skeptical of any argument without data backing it, when it comes to this outage, people are assuming everything written here affects everyone.

Perhaps some of the issues are localized? Perhaps it's even user error (it happens, you know?). But because a small amount of HN users say "it's everywhere!" then suddenly people reach for their pitchforks.

Sometimes we just don't have all the information.

Re: Google Kubernetes Engine's third consecutive day of service disruption

#352

Earlier quoted context omitted.

GCP has a few features that set it apart from other cloud providers. GKE is head and shoulders above the other offerings from AWS and Azure. GCP can be a fair bit cheaper than AWS and Azure for certain workloads. Raw compute/memory is about the same. Storage can make a big difference. GCP persistent SSD costs a bit more than AWS GP2 with much better performance and way cheaper than IO2. Local SSD is also way, way che…

I have a more detailed post above, but if you are large enough, you’re not paying the listed price for AWS. But even if you are, prices change all of the time. From a completely selfish standpoint, is the price difference worth the cost to bet your reputation on if you are the one that made the final decision? Even if statistically the same could happen with AWS, no one would blame you for choosing AWS. However, I co…

To be fair, prices are stable across the providers and Google cloud is very competitive.

It's not like 5 years ago when everyone was ramping up their offerings with a yearly price drop and a new generation.

Re: Google Kubernetes Engine's third consecutive day of service disruption

#353

Earlier quoted context omitted.

I think you are downplaying minimal

I don't think OP was intending "minimal" to mean it would be easy to get to the stage where it's possible, just that once you've got all your infrastructure-as-code stuff set up correctly, you ought to be able to just be pressing buttons / running scripts and have your infrastructure up and running in another cloud provider. Even when working in small companies with small infrastructure, I've kept recreation of infra…

The infrastructure has always been the easy part, as long as the company is willing to pay for multiple datacenters.

Then you realize a lot of software and databases can only run from a single instance, zero support for multi regions, and you're not gonna to rewrite everything and resiliency just can't happen.

Re: Google Kubernetes Engine's third consecutive day of service disruption

#354

Earlier quoted context omitted.

As someone who works for Government and Enterprise - all I care about sometimes is how a company behaves when everything goes wrong. The issue with outages for the Government organizations I have dealt with is rarely the outage itself - but strong communication about what is occurring and realistic approximate ETAs, or options around mitigation. Being able to tell the Directors/Senior managers that issues have been "…

As a government or large enterprise, you should get a support contract with the provider and have a dedicated support to contact. Don't get it wrong. AWS is the exact same thing as Google. All you will is log a ticket and receive an automated ack by the next day.

You are incorrect about aws. If your pay for business support, and something is happening to your production environment, they are on a call with you in less than an hour.

Re: Google Kubernetes Engine's third consecutive day of service disruption

#355

Earlier quoted context omitted.

Most small companies on AWS with revenue outsource support to an MSP.

Any stats / surveys on this? Most smaller (< 50 employees) b2c enterprise-style SaaS companies I'm anecdotally familiar with barely have enough funds to hire ops engineers let alone outsource to MSPs (even if the MSP practices labor rate arbitrage). A lot of the criteria I'd wager may be conditions around pre-revenue status and funding conditions moreso than headcount. I'm trying to understand just how biased my own…

I don’t know about pre revenue companies, but, if you have to be secure and compliant with regulations from day one, you’re not going to take a chance with having a bunch of devs setting up your infrastructure. My bias probably comes from working mostly with companies in highly regulated fields.

Besides, I’m assuming that the cost savings a small company can get from being billed under a much larger organization account would make up for it. That and having cheap shared netops support.

Re: Google Kubernetes Engine's third consecutive day of service disruption

#356

Earlier quoted context omitted.

You're right in terms of breadth officially covered. But if you look at the features where they both officially have support, there are many examples where the GCP version is more reliable and usable than the AWS version. Even GKE is an example of this, despite the outage in node pool creation that we're discussing here. Way better than EKS. (Disclosure: I worked for Google, including GCP, for a few years ending in 2…

I think you're going to have to back up a claim like this with some facts. GKE being the exception, since it was launched a couple years before EKS. AWS clearly has way more services, and the features are way deeper than GCP. Just compare virtual machines and managed databases, AWS has about 2-3x more types of VMs (VMs with more than 4TB of RAM, FPGAs, AMD Epyc, etc.), and in databases, more than just MySQL and Postg…

Each platform has features the other platform doesn't, even though AWS has more.

Some of GCP's unique compelling features include live VM migration that makes it less relevant when a host has to reboot, the new life that has recently been put into Google App Engine (both flexible environment and the second generation standard environment runtimes), the global load balancer with a single IP and no pre-warming, and Cloud Spanner.

In terms of feature coverage breadth I started my previous comment by agreeing that AWS was ahead, and I still reaffirm that. But if you randomly select a feature that they both have to a level which purports to meet a given customer requirement, the GCP offering will frequently have advantages over the AWS equivalent.

Examples besides GKE: BigQuery is better regarded than Amazon Redshift, with less maintenance hassle. And EC2 instance, disk, and network performance is way more variable than GCE which generally delivers what it promises.

One bit of praise for AWS: when Amazon does document something, the doc is easier to find and understand, and one is less likely to find something out of date in a way that doesn't work. But GCP is more likely to have documented the thing in the first place, especially in the case of system-imposed limits.

To be clear, I want there to be three or four competitive and widely used cloud options. I just think GCP is now often the best of the major players in the cases where its scope meets customer needs.

Re: Google Kubernetes Engine's third consecutive day of service disruption

#357

Earlier quoted context omitted.

Any stats / surveys on this? Most smaller (< 50 employees) b2c enterprise-style SaaS companies I'm anecdotally familiar with barely have enough funds to hire ops engineers let alone outsource to MSPs (even if the MSP practices labor rate arbitrage). A lot of the criteria I'd wager may be conditions around pre-revenue status and funding conditions moreso than headcount. I'm trying to understand just how biased my own…

Small companies put every developer on call and here comes the support coverage. Of course, that doesn't make them knowledgeable to run stable infrastructure and they will move on as soon as they realize they are being abused to work overnight and week end.

I am developer and I have the knowledge to run stable infrastructure at least on a small company SAAS scale. But, I wouldn’t go near a company that expected developers to be on call for netops work.

A company that doesn’t want the overhead of an MSP which in my experience is less than the cost of a full time Dev is not a company I’m going to work for. It would tell me a lot about thier mentality.

Re: Google Kubernetes Engine's third consecutive day of service disruption

#358

Earlier quoted context omitted.

> instead they keep sending tickets back for "more info" Isn't that the case with basically every support request, no matter the company or severity? The first couple of emails from 1st & even 2nd level support are mostly about answering the same questions about the environment over and over again. We've had this ping-pong situation with production outages (which we eventually analysed and worked around by ourselves)…

I've definitely had interactions with smaller companies where you can effectively bypass first and second line by demonstrating you know what you're doing, mostly just saying the right things for them to accept that you've done basic troubleshooting steps already and really do need to talk to someone beyond that point.

Yes, same experience here, support at smaller companies can be more dedicated when talking to "knowledgeable" customers. It's generally easier to get to their 3rd level, sometimes just because there is no 1st or 2nd level at all. But at "big" enterprises - not so much.

Re: Google Kubernetes Engine's third consecutive day of service disruption

#359
post #207

Earlier quoted context omitted.

That’s not how Terraform works. Each provisioner has separate syntax depending on the cloud provider. The template for AWS wouldn’t work on GCP or Azure.

Correct. However written properly you get 90% of the way there. Disaster recovery by switching to another provider is simple when minimal centos/rhel images are used.

If all you are doing with your cloud provider are a bunch of VMs, you’re wasting money. There are much cheaper and simpler ways just to host a bunch of VMs than to use one of the major cloud providers.

Are you not using any of thier managed services and are you maintaining your own on VMs? If so, you have the worse of both worlds. You’re spending more on hosting and you’re not saving money on letting someone else do the “undifferentiated heavy lifting”.

Post reply on HN