Google Cloud networking issues in us-east1
151–160 of 341 posts
Re: Google Cloud networking issues in us-east1
#152Earlier quoted context omitted.
Terrance here from Google Cloud Support. There are only 3 things I can say about this situation. 1) These issues are currently unrelated. 2) We learn a lot from these situations. 3) A lot of these types of issues can be mitigated by running in more then 1 region. I really cant promise that today's situations will never happen again. There are a lot of moving pieces in our system and sometimes there are things outside…
Thanks for the reply Terrance. But isn't it more expensive to run in more than one region?
For some customer it is the right thing for other customers it may not be the right thing.
Every provider will have failures. So the question mostly boils down to does paying for more then 1 region cost more or less then paying for the the lost productivity or revenue of an outage like this.
From some places the most costly things they spend money on is employees. If your whole company comes to a stop for even 1 hour. It may cost more then the engineering effort for multi zone, multi region or multi cloud for your critical environments.
Re: Google Cloud networking issues in us-east1
#153Hacker News: The real status page and help desk for the internet. Do companies realize how absurd this is? ETA: It seems someone at Google had a change of heart, and most of what boulos posted in this thread has been added as updates to the official google status page. Better late than never, I guess, especially if this is the start of a trend in outage reporting.
I mostly responded because there was confusion downthread (and in the title) about being “down”. During an outage is a tricky time for comms, so short corrections are best until a full postmortem can be done.
Re: Google Cloud networking issues in us-east1
#154Earlier quoted context omitted.
Google seems to be more forthcoming with their issues. We have seen incidents in AWS where the status never got updated, but support confirmed issues.
Show me a GCP post-mortem that's as detailed and proactive about future improvement as https://status.aws.amazon.com/s3-20080720.html Their last one was laughable in it's lack of self-awareness.
Can you explain what's better about the AWS one? They both do, approximately, the same thing: provide a few paragraphs of background, approximately one paragraph describing the actual issue, and a few paragraphs describing concrete followups. The AWS one has more timestamps.
You aren't confusing this[0] with the postmortem, are you?
[0]: https://cloud.google.com/blog/topics/inside-google-cloud/an-...
Re: Google Cloud networking issues in us-east1
#155Pretty sure I've read before that us-east1 is one of the older Google data centers presumably with older equipment
I think you’re thinking of AWS’s us-east-1 in Virginia. I don’t recall when us-east1 for us was constructed, but this wasn’t any sort of “old equipment” issue. Even there, while your experience may vary, AWS certainly has both old and new equipment.
Re: Google Cloud networking issues in us-east1
#156It's been down for 4 hours and it's just now being posted on HN? Is it intermittent?
There were (and continue to be) connectivity issues due to a subset of the fiber links having trouble. But that’s different from being “down”, it’s “just” an outage. We won’t declare the outage over until the impact is minimal.
Re: Google Cloud networking issues in us-east1
#157Why so many problems at Google lately? Calendar down two weeks ago[0], and Google Cloud had a larger outage a month ago[1] [0]: https://news.ycombinator.com/item?id=20213092 [1]: https://news.ycombinator.com/item?id=20077421
Terrance here from Google Cloud Support. There are only 3 things I can say about this situation. 1) These issues are currently unrelated. 2) We learn a lot from these situations. 3) A lot of these types of issues can be mitigated by running in more then 1 region. I really cant promise that today's situations will never happen again. There are a lot of moving pieces in our system and sometimes there are things outside…
Re: Google Cloud networking issues in us-east1
#158Earlier quoted context omitted.
Terrance here from Google Cloud Support. There are only 3 things I can say about this situation. 1) These issues are currently unrelated. 2) We learn a lot from these situations. 3) A lot of these types of issues can be mitigated by running in more then 1 region. I really cant promise that today's situations will never happen again. There are a lot of moving pieces in our system and sometimes there are things outside…
“You should be using more than 1 region” could also be “you should be using more than one provider”, no?
And, not to be snarky, but many of the other responses that are along the lines of "It's not really that difficult to run in multiple clouds" - let's just say I have trouble believing these commenters have real world experience actually doing this. I'm not saying it's impossible, but it is extremely difficult for any system of reasonable complexity with a dev team of, say, 10 or more people.
And, if you can stomach the cost, you do give up the ability to really use any of the proprietary (and often times awesome) functionality of a particular provider, which can put your dev velocity at a big disadvantage.
Re: Google Cloud networking issues in us-east1
#159Hacker News: The real status page and help desk for the internet. Do companies realize how absurd this is? ETA: It seems someone at Google had a change of heart, and most of what boulos posted in this thread has been added as updates to the official google status page. Better late than never, I guess, especially if this is the start of a trend in outage reporting.
Re: Google Cloud networking issues in us-east1
#160Hacker News: The real status page and help desk for the internet. Do companies realize how absurd this is? ETA: It seems someone at Google had a change of heart, and most of what boulos posted in this thread has been added as updates to the official google status page. Better late than never, I guess, especially if this is the start of a trend in outage reporting.
seriously, they've got a text field on the official status page, why not put the text boulos posted here in that instead of the meaningless text they've got there?