Live data from Hacker News

Google Kubernetes Engine's third consecutive day of service disruption

status.cloud.google.com

341–350 of 419 posts

Re: Google Kubernetes Engine's third consecutive day of service disruption

#341

Earlier quoted context omitted.

Most small companies on AWS with revenue outsource support to an MSP.

Any stats / surveys on this? Most smaller (< 50 employees) b2c enterprise-style SaaS companies I'm anecdotally familiar with barely have enough funds to hire ops engineers let alone outsource to MSPs (even if the MSP practices labor rate arbitrage). A lot of the criteria I'd wager may be conditions around pre-revenue status and funding conditions moreso than headcount. I'm trying to understand just how biased my own…

Small companies put every developer on call and here comes the support coverage.

Of course, that doesn't make them knowledgeable to run stable infrastructure and they will move on as soon as they realize they are being abused to work overnight and week end.

Re: Google Kubernetes Engine's third consecutive day of service disruption

#342
post #234

Earlier quoted context omitted.

> When an entire region is down what I have noticed is that all things are fucked globally on aws. Do you have an example on this?

On 17 October, there was a multi-AZ network failure at us-east-1. It only lasted 3m35s, but it was enough that our customers were calling about our site being down.

That's still just one region, unless you were also hosted outside us-east-1.

Re: Google Kubernetes Engine's third consecutive day of service disruption

#343
post #303

Earlier quoted context omitted.

You've managed to roll the myth of 10x engineers, start-up geniuses, nostalgia and gut feelings into one message of very dubious veracity. At the same time you ignored the massive complexity and size of Google compared to what they were at the beginning. This is voodoo organisational analysis.

10x engineers are real. Start-up geniuses are real. The large majority of people have their heads up their asses. Wake up and smell the coffee

You can't find and hire geniuses for every component. Systems should scale with the average engineer in mind (so should code). We have all smelt the coffee and it smells even better when your team is efficient and well rested.

Re: Google Kubernetes Engine's third consecutive day of service disruption

#344
post #233

Earlier quoted context omitted.

Just grabbed first article. Example: In this case capitalone went down. I don’t work at capitalone - but I imagine they had their data copied across every region 30 times. https://www.geekwire.com/2018/widespread-outage-amazon-web-s...

I think you're much too optimistic about capitalone. They probably had a single point of failure, possibly one they didn't realize they had.

CapitalOne is one of the few financial markets firms with open source cloud projects on github. I respect their tech org for that.

Re: Google Kubernetes Engine's third consecutive day of service disruption

#345
post #253

Earlier quoted context omitted.

Let me know the next time you hear about the CIO of a fortune 500 asking his technology practitioners to validate what he read in Gartner and heard from Diane Greene.

My advice would be to find opportunities to get paid to tell people the right answer, not to implement the wrong answer against your better judgement. Hot job market right now, more jobs than talent, all that jazz. If you're stuck implementing a suboptimal solution, that's not your fault, and not the intent of my above comment.

Lots of wisdom in this comment

Re: Google Kubernetes Engine's third consecutive day of service disruption

#346

Earlier quoted context omitted.

When everything works, GCP is the best. Stable, fast, simple, reliable. When things stop working, GCP is the worst. Slow communications and they require way too much work before escalating issues or attempting to find a solution. They already have the tools and access so most issues should take minutes for them to gather diagnostics, but instead they keep sending tickets back for "more info", inevitably followed by a…

As someone who works for Government and Enterprise - all I care about sometimes is how a company behaves when everything goes wrong. The issue with outages for the Government organizations I have dealt with is rarely the outage itself - but strong communication about what is occurring and realistic approximate ETAs, or options around mitigation. Being able to tell the Directors/Senior managers that issues have been "…

As a government or large enterprise, you should get a support contract with the provider and have a dedicated support to contact.

Don't get it wrong. AWS is the exact same thing as Google. All you will is log a ticket and receive an automated ack by the next day.

Re: Google Kubernetes Engine's third consecutive day of service disruption

#347
post #86
post #45

Earlier quoted context omitted.

Yeah, AWS has never had a global outage, Google has had 2 now.

I'm not sure where you get this idea, but AWS has definitely had global outages. 1.5 years ago there were massive global issues with S3, causing even their own status dashboard to malfunction. edit: I stand corrected. Apparently the S3 outage wasn't global, though its effects were. Meanwhile, this outage has only really been noticeable to ops teams, since it doesn't affect existing nodes or anything outside GKE. It's…

I worked at a company that was highly dependent on us-east at the time. We survived. This is very different

Re: Google Kubernetes Engine's third consecutive day of service disruption

#348

Earlier quoted context omitted.

When everything works, GCP is the best. Stable, fast, simple, reliable. When things stop working, GCP is the worst. Slow communications and they require way too much work before escalating issues or attempting to find a solution. They already have the tools and access so most issues should take minutes for them to gather diagnostics, but instead they keep sending tickets back for "more info", inevitably followed by a…

> instead they keep sending tickets back for "more info" Isn't that the case with basically every support request, no matter the company or severity? The first couple of emails from 1st & even 2nd level support are mostly about answering the same questions about the environment over and over again. We've had this ping-pong situation with production outages (which we eventually analysed and worked around by ourselves)…

I've definitely had interactions with smaller companies where you can effectively bypass first and second line by demonstrating you know what you're doing, mostly just saying the right things for them to accept that you've done basic troubleshooting steps already and really do need to talk to someone beyond that point.

Re: Google Kubernetes Engine's third consecutive day of service disruption

#349
post #57

Earlier quoted context omitted.

It's weekend, why wouldn't you take it off? It's just silly software.

More to the point, why would you depend on Google for any critical infrastructure after this?

Then in a few months from now, people will be saying the same things after an AWS or Azure outage.

Am I the only one that finds this slightly humorous?

Re: Google Kubernetes Engine's third consecutive day of service disruption

#350
post #284

Earlier quoted context omitted.

Google Maps has never been subjected to that policy, unlike GCP services. These org chart divisions are real but only clear to Googlers, Xooglers (I'm in this category), and people who pay extremely close attention. The fact that they're all Google makes reputation damage bleed across meaningfully different parts of what's in truth now a conglomerate under the umbrella name Google.

Except all the Google maps setup and API keys are generated from the gcp UI and the billing happens on the cloud platform as well. While maps didn't start as a gcp product, they seem to have rolled it in to gcp fully.

Not fully. Really what happened is they did a re-org that gave them Google Cloud as an umbrella brand including GCP, Google Maps Platform (this new version of Google Maps as a commercial service), Chrome, Android, G Suite...

The bit of Maps Platform integration for management of the billing and API layer was called out in the announcement blog as an integration with the console specifically, and the docs and other branding around Maps Platform remain distinct from GCP still in excessively subtle ways that Googlers pay more attention to than everyone else, like hosting the docs on developers.google.com instead of cloud.google.com and having Platform in its name separately from Cloud Platform.

This stuff makes sense to Googlers not only because of the org chart but also because Google has a pretty unified API layer technology and because Google put in a lot of work to unify billing tech & management. Reusing that is efficient but not always clear.

But you're right to be confused. Their branding is a mess and always has been. This is the same company that thought Google Play Books makes sense as a product name.

Google's product / PR / comms / exec people are very bad at understanding how external people who don't know Google's org chart and internal tech will perceive these things, or at least bad at prioritizing those concerns.

They live and breathe their corporate internals too much to realize this. Some Google engineers and tech writers realize the confusion but pick other battles to fight instead (like making good quality products).

They do at least document which services are subjected to the GCP Deprecation Policy (Maps is not there): https://cloud.google.com/terms/deprecation

As for what products are actually part of GCP, it's the parts of this page that aren't an external partner's brand name, aren't called out separately like G Suite or Cloud Identity or Cloud Search, and aren't purely open source projects like Knative and Istio (as opposed to the productized versions within GCP), with the caveat that the level so far of integration into GCP of Google acquisitions like Apigee, Firebase, and Stackdriver varies depending on per-company specifics: https://cloud.google.com/products/

G Suite and Cloud Identity accounts can be used with GCP, just like any other Google accounts. They are part of Google Cloud but not Google Cloud Platform.

Hope I waded through the mess correctly for you. :)

Post reply on HN