Google Cloud networking issues in us-east1
201–210 of 341 posts
Re: Google Cloud networking issues in us-east1
#202Disclosure: I work on Google Cloud (but I'm not in SRE, oncall, etc.). As the updates to [1] say, we're working to resolve a networking issue. The Region isn't (and wasn't) "down", but obviously network latency spiking up for external connectivity is bad. We are currently experiencing an issue with a subset of the fiber paths that supply the region. We're working on getting that restored. In the meantime, we've remov…
Tangential question: does Google allow employees, not directly tasked with it, to represent the company online as they wish? Most companies I know of have a strict ‘do not speak for the company’ policy.
In my case, Cloud PR knows me, but I also knowingly risk my job (I clearly believe I have good enough judgment in what I post). If Urs and Ben think I should be fired, I'm okay with that, as it would represent a significant enough difference in opinion, that I wouldn't want to continue working here anyway.
Finally, for what it's worth, I have been reported before for "leaking internal secrets" here on HN! It turned out to be a totally hilarious discussion with the person tasked with questioning me. Still not fired, gotta try harder :).
Re: Google Cloud networking issues in us-east1
#203Earlier quoted context omitted.
Since we use GCP Global LBs I presume that "draining the Google.com traffic" also meant that you're diverting all global LB traffic, which is what we see. The second incident (the OP's link) indicates that but at first it was very confusing to a customer when the first issue was marked as resolved but we still saw no traffic being sent to us-east1 via our global LBs. If that makes sense.
This part was somewhat nuanced, so I wasn’t sure to post it: yes, if you are using GCLB, and have more than 1 healthy Region, we will also rebalance to avoid us-east- for now (though not so statically as that sounds, mumble mumble). Edit: added this to the top level comment so more folks see it.
Re: Google Cloud networking issues in us-east1
#204Earlier quoted context omitted.
How can Stack Overflow run on a single server? Do you mean single cluster?
As of 2016, Stack Overflow ran on dozens of servers in two data centers. https://nickcraver.com/blog/2016/03/29/stack-overflow-the-ha...
Re: Google Cloud networking issues in us-east1
#205Earlier quoted context omitted.
It's not independent fiber links if they use the same tube to get into the building...just ask any backhoe operator.
my brother-in-law's construction company actually did just that. ground wasn't properly marked and the fiber got cut, multiple links
Re: Google Cloud networking issues in us-east1
#206Does anybody else feel like there have been a lot of outages in recent months? And I don't mean Google -- I mean lots of others too (I seem to recall CloudFlare, Facebook, etc.)... are they really increasing or are we just hearing more about them? Seems a bit odd.
Re: Google Cloud networking issues in us-east1
#207Earlier quoted context omitted.
Terrance here from Google Cloud Support. There are only 3 things I can say about this situation. 1) These issues are currently unrelated. 2) We learn a lot from these situations. 3) A lot of these types of issues can be mitigated by running in more then 1 region. I really cant promise that today's situations will never happen again. There are a lot of moving pieces in our system and sometimes there are things outside…
“You should be using more than 1 region” could also be “you should be using more than one provider”, no?
Multiple regions, as long as your provider offers all of the services, you can have a carbon copy. Much easier.
It depends on your needs, your architecture, your risk tolerance, etc. I think for most people "Use multiple regions" is the answer that strikes the correct balance. It probably isn't the correct answer for everyone.
Re: Google Cloud networking issues in us-east1
#208Earlier quoted context omitted.
Well, sure, if you hate your devops team and you want to make sure they can’t use any of the proprietary functionality of either provider. At which point, if you want to be managing a fleet of vanilla Linux boxes yourself, why use a cloud provider at all?
Why would you want to lock into a cloud provider? You're losing a lot of operational flexibility for less devops and sysaadmin work. You are really limiting your tech stack by using standardized things like Jenkins, Docker, K8, mqtt, kafka.
"Outsourcing" those functions to cloud services can be big win for a small team. Like all engineering, it's a trade off.
Re: Google Cloud networking issues in us-east1
#209Does anybody else feel like there have been a lot of outages in recent months? And I don't mean Google -- I mean lots of others too (I seem to recall CloudFlare, Facebook, etc.)... are they really increasing or are we just hearing more about them? Seems a bit odd.
The more "the cloud" replaces many, many servers at lots of different places, the more the outages (which once happened all the time, but to many different organizations at different times) will become big enough to notice.
So, yeah, not just your imagination.
Re: Google Cloud networking issues in us-east1
#210Does anybody else feel like there have been a lot of outages in recent months? And I don't mean Google -- I mean lots of others too (I seem to recall CloudFlare, Facebook, etc.)... are they really increasing or are we just hearing more about them? Seems a bit odd.
It's almost as if we had made an overly complicated system with too much "efficiency" and thus not enough redundancy, centralizing on too few pieces of what used to be a quite widely dispersed system. The more "the cloud" replaces many, many servers at lots of different places, the more the outages (which once happened all the time, but to many different organizations at different times) will become big enough to not…
This is just for the last few months...?