Live data from Hacker News

Google Kubernetes Engine's third consecutive day of service disruption

status.cloud.google.com

321–330 of 419 posts

Re: Google Kubernetes Engine's third consecutive day of service disruption

#321
post #5
post #2

Status page is inaccurate as issues doesn't only affect the web UI, the same operations are not functioning via the CLI.

Its kinda strange that HN seems to be the most effective way to give feedback to Google Cloud :/

Most transparent, at least.

Re: Google Kubernetes Engine's third consecutive day of service disruption

#322

Earlier quoted context omitted.

When everything works, GCP is the best. Stable, fast, simple, reliable. When things stop working, GCP is the worst. Slow communications and they require way too much work before escalating issues or attempting to find a solution. They already have the tools and access so most issues should take minutes for them to gather diagnostics, but instead they keep sending tickets back for "more info", inevitably followed by a…

As someone who works for Government and Enterprise - all I care about sometimes is how a company behaves when everything goes wrong. The issue with outages for the Government organizations I have dealt with is rarely the outage itself - but strong communication about what is occurring and realistic approximate ETAs, or options around mitigation. Being able to tell the Directors/Senior managers that issues have been "…

Very similar thing at our office. Considering the scale of which we run things, any outage could be a potential loss of millions _every minute_.

Sure, we use support tickets with vendors for small things. Console button bugging out, etc. But for large incidents, every vendor has a representative within an hour driving distance and will be called into a room with our engineers to fix the problem. This kind of outage, with zero communication, means the dropping of a contract.

Communication is critical for trust, especially if we're running a business off it.

Re: Google Kubernetes Engine's third consecutive day of service disruption

#323

We had an issue a few weeks ago where the google front-end servers were mangling responses from Pub/Sub and returning 502 responses, making the service completely unusable and knocking over a number of things we have running in production. Despite paying for enterprise support and having in a P1 ticket, we had to spend Friday to Sunday gathering evidence to prove to the support staff that there was indeed a problem,…

I just took a screenshot of this and any time somebody asks my why AWS i will forward them. Thanks!

Re: Google Kubernetes Engine's third consecutive day of service disruption

#324

Earlier quoted context omitted.

When everything works, GCP is the best. Stable, fast, simple, reliable. When things stop working, GCP is the worst. Slow communications and they require way too much work before escalating issues or attempting to find a solution. They already have the tools and access so most issues should take minutes for them to gather diagnostics, but instead they keep sending tickets back for "more info", inevitably followed by a…

To say "when it works it's stable and reliable" implies that it is neither...

60% of the time, it works every time...

Re: Google Kubernetes Engine's third consecutive day of service disruption

#325
post #312

Earlier quoted context omitted.

When everything works, GCP is the best. Stable, fast, simple, reliable. When things stop working, GCP is the worst. Slow communications and they require way too much work before escalating issues or attempting to find a solution. They already have the tools and access so most issues should take minutes for them to gather diagnostics, but instead they keep sending tickets back for "more info", inevitably followed by a…

"Support costs" calculation often doesn't include the costs of not having support. When I worked at GoDaddy, there were around 2/3 of the company was customer support. At the current company I'm at, a cryptocurrency exchange, our support agents frequently hear they prefer our service over others because of our fast support response times (crypto exchanges are notorious for really poor support). All of my interactions…

Google hasn't learned this lesson.

They have though; they've just drawn the conclusion that they'd rather put massive amounts of effort in to building services that users can use without needing support. This approach works well once the problems have been ironed out, but it's horrible until that's the case. Google's mature products like Ads, Docs, GMail, etc are amazing. Their new products ... aren't.

Re: Google Kubernetes Engine's third consecutive day of service disruption

#326

Earlier quoted context omitted.

We are GCP customers for the last couple of years. We use other cloud platforms(AWS, IBM, Oracle, OrionVM) too. We don't use GKE but use rancher/kubernetes combo on their standard platform. So far GCP is the best, hands down in terms of stability. We never had a single outage or maintenance downtime notification till now. We are power users but our monitoring didn't pick any anomaly so i don't think this issue had ra…

I don't understand why someone would choose to deploy anything mission critical without having an support contract with the ISP, the manufacturer of the the software etc.

Or why chose an error prone technology?

Re: Google Kubernetes Engine's third consecutive day of service disruption

#327

Earlier quoted context omitted.

This is why GCP has no hope of ever taking significant market share from AWS. Google thinks they can treat their cloud customers like they treat users of their free services. Customer support and communication are essential.

As if something like this has never happened to AWS?

Not effecting all regions no.

Re: Google Kubernetes Engine's third consecutive day of service disruption

#328
post #312

Earlier quoted context omitted.

"Support costs" calculation often doesn't include the costs of not having support. When I worked at GoDaddy, there were around 2/3 of the company was customer support. At the current company I'm at, a cryptocurrency exchange, our support agents frequently hear they prefer our service over others because of our fast support response times (crypto exchanges are notorious for really poor support). All of my interactions…

Google hasn't learned this lesson. They have though; they've just drawn the conclusion that they'd rather put massive amounts of effort in to building services that users can use without needing support. This approach works well once the problems have been ironed out, but it's horrible until that's the case. Google's mature products like Ads, Docs, GMail, etc are amazing. Their new products ... aren't.

There's a big difference between SaaS applications and compute infrastructure for your business.

Google Ads and such also have a terrible support reputation, even with clients spending 8 figures.

Re: Google Kubernetes Engine's third consecutive day of service disruption

#329

Earlier quoted context omitted.

The preceding line is the imprtant part. Which is essentially "Subject to the deprecation policy [which says that Google will give at least 1 year notice before cancelling services], Google may discontinue..." In other words, at any time, google can give you a years notice. (I work at Google, but am not a lawyer and this isn't official in any capacity). Please don't selectively quote things out of context to give a m…

But what do these things mean? > commercially reasonable > substantial economic or material technical burden Is one engineer working on an old service to keep it alive commercially reasonable or a substantial burden? I don't know. Do you? In practice this policy lets them shut off anything they want any time they want. Again it's their playground they can do what they want unless they signed a contract saying they'd…

By the way, that's probably more than $100k/year if he does. Probably triple that with the overhead of the office space, HR and middle management.

Re: Google Kubernetes Engine's third consecutive day of service disruption

#330

"third consecutive day of service disruption" is not an accurate statement? Latest update was Nov 11 saying things resolved on Nov 9. https://status.cloud.google.com/incident/container-engine/18...

If all nodes in GKE clusters were down for 3 days, I would consider this newsworthy and shocking. This... is not. Come on, people.
Post reply on HN