Live data from Hacker News

Google Kubernetes Engine's third consecutive day of service disruption

status.cloud.google.com

271–280 of 419 posts

Re: Google Kubernetes Engine's third consecutive day of service disruption

#271
post #121

Earlier quoted context omitted.

They work for Google so obviously they are much smarter than you. If theres a problem its probably the customers fault. /sarcasm

I was so mad to read that until you said /sarcasm :p That being said I really do think there is a difference between who is working at google today and the google we all fell in love with pre-2008. I am sure there are a amazing people still working at google, but nowhere near like it was. The way I like to think about google is that some amazing people mad ea awesome train that builds tracks in front of it -- you can…

Even though you have troubleshooting skills and other skills that can help Google, Google assumes that these skills are derivative of what they are looking for in interview. So, they are looking for those primary skills.

Re: Google Kubernetes Engine's third consecutive day of service disruption

#272
post #232

Earlier quoted context omitted.

The most recent example that I can think of when something went down globally was Route 53 - the one service that AWS promises 100% up time for.

Citation needed?

See the reply to my original post. It wasn’t an actually Route 53 “outage”.

Re: Google Kubernetes Engine's third consecutive day of service disruption

#273

We had an issue a few weeks ago where the google front-end servers were mangling responses from Pub/Sub and returning 502 responses, making the service completely unusable and knocking over a number of things we have running in production. Despite paying for enterprise support and having in a P1 ticket, we had to spend Friday to Sunday gathering evidence to prove to the support staff that there was indeed a problem,…

Again, because their support reps don't believe there's a problem.

Perhaps if you explained it on a whiteboard...

Re: Google Kubernetes Engine's third consecutive day of service disruption

#274
post #18

Earlier quoted context omitted.

Same here. We have spent 2 days trying to create instances and migrate images just to figure out later they can't start. Right when I convinced our project to get migrated from AWS...

Same question... why would you do that? AWS is super stable most of the time. I have been running k8s over EC2 (not eks) for a year and works like a charm. I've even run experiments using spot instances and it's pretty good (no guarantee there of course).

We did. Our customer had a bucket shared to a role they created, we spent a week back and forth trying to mount such bucket in our instance using fuse. Mounting with gcsfuse took 5 minutes (although no role used, so perhaps an unfair comparison). In general, I found Gcloud a lot easier to work with.

Re: Google Kubernetes Engine's third consecutive day of service disruption

#275
post #66

Earlier quoted context omitted.

> I’m sorry to say this, but it is the equivalent of going bankrupt from a trust perspective. It's the opposite really: the expectation that service providers have no unexpected downtime is unrealistic, and it's strange this idea persists.

(disclaimer: I work for another cloud provider) I agree, in general, outages are almost inevitable, but global outages shouldn't occur. It suggests at least a couple of things: 1) Bad software deployments, without proper validation. A message elsewhere in this post on HN suggest that problems have been occurring for at least 5 days, which makes me think this is the most likely situation. If this is the case, presumab…

When I was at google, the big outages were almost always bad routing to the service - it was never that the service couldn't handle the load, and bad service instances were kind of hidden. my service did have some problems on new releases, but because we had multiple instances we could just redirect traffic to the instances we hadn't updated, so they stayed up.

Re: Google Kubernetes Engine's third consecutive day of service disruption

#276

Earlier quoted context omitted.

This is why GCP has no hope of ever taking significant market share from AWS. Google thinks they can treat their cloud customers like they treat users of their free services. Customer support and communication are essential.

As if something like this has never happened to AWS?

"like this" -- a failure of the service, or a failure of communication and customer support?

Re: Google Kubernetes Engine's third consecutive day of service disruption

#277

Earlier quoted context omitted.

As if something like this has never happened to AWS?

"like this" -- a failure of the service, or a failure of communication and customer support?

Both, I suppose.

Re: Google Kubernetes Engine's third consecutive day of service disruption

#278

I am currently evaluating GCP for two separate projects. I want to see if I understand this correctly: 1) For three whole days, it was questionable whether or not a user would be able to launch a node pool (according to the official blog statement). It was also questionable whether a user would be able to launch a simple compute instance (according to statements here on HN). 2) This issue was global in scope, affecti…

When everything works, GCP is the best. Stable, fast, simple, reliable.

When things stop working, GCP is the worst. Slow communications and they require way too much work before escalating issues or attempting to find a solution.

They already have the tools and access so most issues should take minutes for them to gather diagnostics, but instead they keep sending tickets back for "more info", inevitably followed by a hand-off to another team in a different time zone. We have spent days trying to convince them there was an issue before, which just seems unacceptable.

I can understand support costs but there should be a test (with all vendors) where I can officially certify that I know what I'm talking about and don't need to go through the "prove its actually a problem" phase every time.

Re: Google Kubernetes Engine's third consecutive day of service disruption

#279

Earlier quoted context omitted.

I don't understand why someone would choose to deploy anything mission critical without having an support contract with the ISP, the manufacturer of the the software etc.

Simple, the cost of an outage is less than the cost of a support contract. Very few things are really mission critical as in they can never go down. Rather they simply have a cost to going down and you can choose to pay that one way or another.

And it's not like having a support contract precludes you from downtime.

Re: Google Kubernetes Engine's third consecutive day of service disruption

#280

Earlier quoted context omitted.

That shouldn't, but they do. S3 goes down [1]. The AWS global console goes down, right after Prime Day outages [2]. Lots of Google Cloud services go down [3, current thread]. Tens of Azure services go down hard [4]. Are software development and release processes improving to mitigate these outages? We don't know. You have to trust the marketing. Will regions ever be fully isolated? We don't know. Will AWS IAM and con…

Look how frequent and detailed Amazon's update logs are in that first Register article. Multiple updates throughout the day going into some detail.

[deleted]
Post reply on HN