Live data from Hacker News

Google Cloud networking issues in us-east1

status.cloud.google.com

161–170 of 341 posts

Re: Google Cloud networking issues in us-east1

#161
post #140

Here's the original issue: https://status.cloud.google.com/incident/cloud-networking/19... Not sure why they closed that one at 9:12 just to open a new one at 10:25. We didn't see any traffic coming to us-east1 during that time period so I would assume the original issue is still the root cause.

Yeah, that happens sometimes based on which team notices, thinks it might be different and then opens an outage. Sorry for the confusion, and yes, the fiber link issue is the root cause. Draining the Google.com traffic presumably resolved the issue for you, though you may still be seeing elevated latency as the updates suggest.

Since we use GCP Global LBs I presume that "draining the Google.com traffic" also meant that you're diverting all global LB traffic, which is what we see. The second incident (the OP's link) indicates that but at first it was very confusing to a customer when the first issue was marked as resolved but we still saw no traffic being sent to us-east1 via our global LBs. If that makes sense.

Re: Google Cloud networking issues in us-east1

#162

Earlier quoted context omitted.

Terrance here from Google Cloud Support. There are only 3 things I can say about this situation. 1) These issues are currently unrelated. 2) We learn a lot from these situations. 3) A lot of these types of issues can be mitigated by running in more then 1 region. I really cant promise that today's situations will never happen again. There are a lot of moving pieces in our system and sometimes there are things outside…

“You should be using more than 1 region” could also be “you should be using more than one provider”, no?

Maybe. If you get a billing issue or get marked as suspicious, you can lose all services with one provider.

Re: Google Cloud networking issues in us-east1

#163

Hacker News: The real status page and help desk for the internet. Do companies realize how absurd this is? ETA: It seems someone at Google had a change of heart, and most of what boulos posted in this thread has been added as updates to the official google status page. Better late than never, I guess, especially if this is the start of a trend in outage reporting.

You seem to have forgotten twitter

We can dream.

Re: Google Cloud networking issues in us-east1

#164

Hacker News: The real status page and help desk for the internet. Do companies realize how absurd this is? ETA: It seems someone at Google had a change of heart, and most of what boulos posted in this thread has been added as updates to the official google status page. Better late than never, I guess, especially if this is the start of a trend in outage reporting.

seriously, they've got a text field on the official status page, why not put the text boulos posted here in that instead of the meaningless text they've got there?

I work for AWS. There is typically a balance that has to be struck when sharing information with customers. I would imagine this goes for most companies, which is why it isn't until a post-mortem that the messaging is fully refined.

Re: Google Cloud networking issues in us-east1

#165
post #140

Earlier quoted context omitted.

Yeah, that happens sometimes based on which team notices, thinks it might be different and then opens an outage. Sorry for the confusion, and yes, the fiber link issue is the root cause. Draining the Google.com traffic presumably resolved the issue for you, though you may still be seeing elevated latency as the updates suggest.

Since we use GCP Global LBs I presume that "draining the Google.com traffic" also meant that you're diverting all global LB traffic, which is what we see. The second incident (the OP's link) indicates that but at first it was very confusing to a customer when the first issue was marked as resolved but we still saw no traffic being sent to us-east1 via our global LBs. If that makes sense.

This part was somewhat nuanced, so I wasn’t sure to post it: yes, if you are using GCLB, and have more than 1 healthy Region, we will also rebalance to avoid us-east- for now (though not so statically as that sounds, mumble mumble).

Edit: added this to the top level comment so more folks see it.

Re: Google Cloud networking issues in us-east1

#166

Earlier quoted context omitted.

Availability of what? I've noticed entire afternoon where it wasn't possible to provision instances of some types, when I was working with AWS daily.

Availability of network access to existing instances. What you're talking about with provisioning capacity is a totally different matter. Provisioning availability is not guaranteed (unless you purchase reserved instances) and there are frequently periods where certain instance types are not available in certain AZs, though they do try to resolve that as fast as practicality allows them to. It really stinks sometimes…

Yeah, you can now provision multiple instance types in an ASG which mitigates this somewhat.

I think people sometimes forget that the cloud isn't magic, and a sudden burst of requests for new instances needs somebody to actually rack up some servers.

Re: Google Cloud networking issues in us-east1

#167

Earlier quoted context omitted.

“You should be using more than 1 region” could also be “you should be using more than one provider”, no?

It's quite common in cloud solution design to design for failure. One of the common assumptions that we hold to is that one region may go down. Other examples: Assume an instance of an app can go down. Assume a VM can go down. Assume a DC can go down. This is not to excuse the downtime in any way.

Do we need a new definition for RAID level?

Redundant Array of independent Data Clouds.

I guess for RAID 5 would I need a min of 3 regions or 3 separate cloud providers.

Re: Google Cloud networking issues in us-east1

#168
post #160

Earlier quoted context omitted.

seriously, they've got a text field on the official status page, why not put the text boulos posted here in that instead of the meaningless text they've got there?

Can you expand on why you find it “meaningless”? As my other comment says, I’m not in SRE and the real people fixing it are trying their best to remediate the problem. I agree that the text I posted (with blessing from SRE!) gives you some more detail, but you can’t do anything differently with it, right? What about the new text do you prefer? (We’re happy to improve!)

Your, even brief, description is interpretable by your clients and some customers - and is actually really informative. It helps estimate the magnitude of the issue, and the types of downstream problems to expect or avoid.

Knowing an astroid took out the entire continent tells you something about the repairability, resources required to fix the problem, and generally provides context for later updates, as opposed to other contexts like a cut fiber line, a burning datacenter or a bad power supply.

Re: Google Cloud networking issues in us-east1

#169
post #13

What's the actual number of 9s for the major cloud services these days? My impression from their PR seems to mismatch the number of outages and issues lately.

AWS EC2 promises 4 9's (4.3 minutes of downtime/month) before their SLA kicks in, but they only give a 10% discount until availability dips below 99% (7.5 hours of downtime/month) when they give a 30% discount. If availability is below 95% (36 hours) in a month, they give a full refund. For an individual instance, they only promise 90% availability.

What a craptactular SLA.

Re: Google Cloud networking issues in us-east1

#170
post #141
post #131

Earlier quoted context omitted.

Do people ever worry that an entire cloud provider may go down, or is that too unlikely of a case?

However much we technical people might salivate at the prospect of designing a multi-cloud solution, for the vast majority of businesses it simply isn't worth the cost / complexity. I'd wager 90-something percent of applications could suffer multi-hour outages without impacting business function to any measurable degree. Plus the fact that without serious investment, you're probably more liable to decrease availabili…

The real trick here, which many people don’t want to look at, is to avoid overly centralizing your workflow.

I can get a lot of work done while Outlook is down. Hell, probably more work done.

If our build server is down I can work for a couple hours (unless we’ve done something very bad). Same for git or our bug database or wiki or or or. When I get stuck on one thing I can swap to something else every couple of hours. And there is always documentation (writing or consuming).

But if some idiot, hypothetically speaking of course, puts most of these services into the same SAN, then we are truly and utterly screwed if there is a hardware failure.

Similarly if you make one giant app that handles your whole business, if that app goes down and there are no manual backups you might as well send everybody home.

I went to get a drink the other day and the place looked funny. They’d tripped a circuit breaker and the whole kitchen lost power. But the registers and the beverage machines were on a separate circuit. And since they sold drinks and food in that order, they stayed open and just apologized a lot. Whoever wired that place knew what they were doing.

Post reply on HN