Live data from Hacker News

Google Cloud networking issues in us-east1

status.cloud.google.com

91–100 of 341 posts

Re: Google Cloud networking issues in us-east1

#91
post #3

Why so many problems at Google lately? Calendar down two weeks ago[0], and Google Cloud had a larger outage a month ago[1] [0]: https://news.ycombinator.com/item?id=20213092 [1]: https://news.ycombinator.com/item?id=20077421

Terrance here from Google Cloud Support. There are only 3 things I can say about this situation. 1) These issues are currently unrelated. 2) We learn a lot from these situations. 3) A lot of these types of issues can be mitigated by running in more then 1 region. I really cant promise that today's situations will never happen again. There are a lot of moving pieces in our system and sometimes there are things outside…

“You should be using more than 1 region” could also be “you should be using more than one provider”, no?

Re: Google Cloud networking issues in us-east1

#92
post #51
post #8

Earlier quoted context omitted.

Availability is hierarchical.

Can you explain that more?

There is no service with 100% availability. You put multiple AZs in one region but nobody was ever pretending that regional failures were impossible, just that single-AZ failures are more common than regional failures. You want high availability, you want multi-regional. Above that you want multi-provider.

The same decisions that make regions fail also makes infra-region traffic cheaper. This is true for all large cloud providers. If you are okay paying more for internal network traffic you can get multiregional. But multi-AZ is still better than single-AZ. Up to you to decide if it’s worth it. For that you need good SLAs and (IMO) support contracts.

Re: Google Cloud networking issues in us-east1

#93
post #77

Kind of related to this, but these types of outages are why I moved from Google Play to Spotify for streaming music. Their infrastructure seems so large that things that should be a standalone service, like streaming music, are bound to be collateral damage when they mess something up on another service. Having everything provided by one company is convenient until it all goes down at the same time and you can't acce…

Spotify is hosted on google cloud: https://www.wired.com/2016/02/spotify-moves-itself-onto-goog...

I think the point OP was trying to make was relating to google services and their dependencies on each other.

Re: Google Cloud networking issues in us-east1

#94

Earlier quoted context omitted.

Terrance here from Google Cloud Support. There are only 3 things I can say about this situation. 1) These issues are currently unrelated. 2) We learn a lot from these situations. 3) A lot of these types of issues can be mitigated by running in more then 1 region. I really cant promise that today's situations will never happen again. There are a lot of moving pieces in our system and sometimes there are things outside…

> There are a lot of moving pieces in our system and sometimes there are things outside of Google's control. Are you implying that the cause of this outage is not Google's fault? If so, can you go into more details about that?

Not him but oftentimes cloud outages can be due to issues with the network connections to the datacenter, or power outages.

Datacenters also sometimes have other single points of failure such as DNS, but those are within the company's control.

https://www.networkworld.com/article/3373646/network-problem...

https://www.datacenterknowledge.com/uptime/equinix-power-out...

Re: Google Cloud networking issues in us-east1

#95

Earlier quoted context omitted.

Terrance here from Google Cloud Support. There are only 3 things I can say about this situation. 1) These issues are currently unrelated. 2) We learn a lot from these situations. 3) A lot of these types of issues can be mitigated by running in more then 1 region. I really cant promise that today's situations will never happen again. There are a lot of moving pieces in our system and sometimes there are things outside…

“You should be using more than 1 region” could also be “you should be using more than one provider”, no?

Yes.

Re: Google Cloud networking issues in us-east1

#96

To whomever commented something like 'laughs in AWS' (comment was removed before I submitted the comment)... please don't... glass house and all that... but I also share the same glass house as you.. I don't want bad luck ... and it's only a fluke that this happened to google in eu-east1 and not AWS in X region and then you (and I) would be having a time of hell! :/

i don't quite follow your logic. something about glass houses and bad luck?

the whole point when something like this happens is for you to ensure that a region going down will not impact you - not to laugh at people that use another cloud or to assume that X is better than Y. That being said, there have been several Google related failures lately that don't help building confidence in the GCP offering - if you're just starting in the cloud space this may actually impact the choices you make when you pick your cloud provider.

Re: Google Cloud networking issues in us-east1

#97
post #13

What's the actual number of 9s for the major cloud services these days? My impression from their PR seems to mismatch the number of outages and issues lately.

AWS EC2 promises 4 9's (4.3 minutes of downtime/month) before their SLA kicks in, but they only give a 10% discount until availability dips below 99% (7.5 hours of downtime/month) when they give a 30% discount. If availability is below 95% (36 hours) in a month, they give a full refund. For an individual instance, they only promise 90% availability.

Availability of what? I've noticed entire afternoon where it wasn't possible to provision instances of some types, when I was working with AWS daily.

Re: Google Cloud networking issues in us-east1

#99
App Engine and Cloud functions were apparently returning error rates of > 30 percent overall between 11 a.m. and 3 p.m., with some projects experiencing a 100 percent error rate. GCS was also experiencing issues for the first half, which was attributed to the networking issues. Google said the networking issues were resolved initially but then stated they were investigating the GAE issues. Those issues were resolved, and the networking issue has been reopened as of 2:35 eastern: https://status.cloud.google.com/incident/cloud-networking/19....

GAE and all other services still show green here, of course: https://status.cloud.google.com/

Re: Google Cloud networking issues in us-east1

#100

Earlier quoted context omitted.

Terrance here from Google Cloud Support. There are only 3 things I can say about this situation. 1) These issues are currently unrelated. 2) We learn a lot from these situations. 3) A lot of these types of issues can be mitigated by running in more then 1 region. I really cant promise that today's situations will never happen again. There are a lot of moving pieces in our system and sometimes there are things outside…

“You should be using more than 1 region” could also be “you should be using more than one provider”, no?

Well, sure, if you hate your devops team and you want to make sure they can’t use any of the proprietary functionality of either provider. At which point, if you want to be managing a fleet of vanilla Linux boxes yourself, why use a cloud provider at all?
Post reply on HN