Live data from Hacker News

Google cloud outage

status.cloud.google.com

11–20 of 39 posts

Re: Google cloud outage

#11
post #9

"We had a router failure in Atlanta". WHAT? You kidding us? Urs Hölzle, technical infrastructure at Google Cloud senior vice president, said, "We're very sorry about that! We had a router failure in Atlanta, which affected traffic routed through that region. Things should be back to normal now. Just to make sure: This wasn't related to traffic levels or any kind of overload, our network is not stressed by COVID-19."

https://twitter.com/uhoelzle/status/1243217659690278912

Re: Google cloud outage

#12
post #4

Heyo Googler here. The problem was a mix between another cloud provider and GCP. Dare I say, there should be little customer impact as of 13:37 PST..... The status dashboard is going to be your best idea on information.

Oh man I had no idea the big cloud providers have dependencies on other clouds like this.

Re: Google cloud outage

#13
post #4

Heyo Googler here. The problem was a mix between another cloud provider and GCP. Dare I say, there should be little customer impact as of 13:37 PST..... The status dashboard is going to be your best idea on information.

Oh man I had no idea the big cloud providers have dependencies on other clouds like this.

Given how much trans-continental/trans-oceanic network cable the major cloud providers own, they almost certainly have special trans-cloud network traffic infrastructure. Especially since so much of "The Cloud" is within a few 10s of square miles in a field in Virginia. I can easily see how one provider could majorly disrupt another provider by accidentally breaking inbound traffic on one of those links.

Re: Google cloud outage

#14
post #9

"We had a router failure in Atlanta". WHAT? You kidding us? Urs Hölzle, technical infrastructure at Google Cloud senior vice president, said, "We're very sorry about that! We had a router failure in Atlanta, which affected traffic routed through that region. Things should be back to normal now. Just to make sure: This wasn't related to traffic levels or any kind of overload, our network is not stressed by COVID-19."

Was it like... a hardware failure? If you serve more than 100 people you probably should have redundant routers. Was it a configuration issue that replicated over to multiple devices at least, I hope?

Not that simple as you sometimes need to manually isolate the faulty hardware and remove it from service.

Re: Google cloud outage

#15
post #9

"We had a router failure in Atlanta". WHAT? You kidding us? Urs Hölzle, technical infrastructure at Google Cloud senior vice president, said, "We're very sorry about that! We had a router failure in Atlanta, which affected traffic routed through that region. Things should be back to normal now. Just to make sure: This wasn't related to traffic levels or any kind of overload, our network is not stressed by COVID-19."

Was it like... a hardware failure? If you serve more than 100 people you probably should have redundant routers. Was it a configuration issue that replicated over to multiple devices at least, I hope?

Have you worked with redundant routers? They certainly reduce the number of outages, but sometimes the hardware (or software) fails in exciting ways that doesn't engage the redundancy, or doesn't engage it properly, and you still get an outage (or you get an outage that wouldn't have happened). Or sometimes, one circuit is out of service for repair or upgrade, and the other circuit is connected to the router that failed. And routing for the AS that travels on that circuit was set not to fallback to transit because the last time that happened, it caused major issues.

I have no specific knowledge of today's events, but this sort of thing happens. You can get the number of incidents down pretty low, but not to zero.

Re: Google cloud outage

#16
post #4

Heyo Googler here. The problem was a mix between another cloud provider and GCP. Dare I say, there should be little customer impact as of 13:37 PST..... The status dashboard is going to be your best idea on information.

[removed]

Re: Google cloud outage

#17
post #9

"We had a router failure in Atlanta". WHAT? You kidding us? Urs Hölzle, technical infrastructure at Google Cloud senior vice president, said, "We're very sorry about that! We had a router failure in Atlanta, which affected traffic routed through that region. Things should be back to normal now. Just to make sure: This wasn't related to traffic levels or any kind of overload, our network is not stressed by COVID-19."

Was it like... a hardware failure? If you serve more than 100 people you probably should have redundant routers. Was it a configuration issue that replicated over to multiple devices at least, I hope?

Networks are harder than everyone thinks. The 2018 CenturyLink outage on the west coast was caused by 1 bad network card that started writing malformed packets.

https://www.geekwire.com/2018/report-huge-centurylink-outage...

Re: Google cloud outage

#18
post #9

"We had a router failure in Atlanta". WHAT? You kidding us? Urs Hölzle, technical infrastructure at Google Cloud senior vice president, said, "We're very sorry about that! We had a router failure in Atlanta, which affected traffic routed through that region. Things should be back to normal now. Just to make sure: This wasn't related to traffic levels or any kind of overload, our network is not stressed by COVID-19."

Was it like... a hardware failure? If you serve more than 100 people you probably should have redundant routers. Was it a configuration issue that replicated over to multiple devices at least, I hope?

Surely the "100 people" metric is too low although I agree at some point (and certainly Google-scale) a redundant router makes sense.

Re: Google cloud outage

#19
post #9

"We had a router failure in Atlanta". WHAT? You kidding us? Urs Hölzle, technical infrastructure at Google Cloud senior vice president, said, "We're very sorry about that! We had a router failure in Atlanta, which affected traffic routed through that region. Things should be back to normal now. Just to make sure: This wasn't related to traffic levels or any kind of overload, our network is not stressed by COVID-19."

Was it like... a hardware failure? If you serve more than 100 people you probably should have redundant routers. Was it a configuration issue that replicated over to multiple devices at least, I hope?

yes, because OBVIOUSLY Google is too stupid to know about redundant routers. /s

https://twitter.com/uhoelzle/status/1243259280410554368

"When routers fail cleanly (say, power out) failover is quick, so you never hear about these. This wasn't such a simple case. We have "many" (not just two) routers in Atlanta so it wasn't an issue of missing redundancy."

Re: Google cloud outage

#20
post #16
post #4

Heyo Googler here. The problem was a mix between another cloud provider and GCP. Dare I say, there should be little customer impact as of 13:37 PST..... The status dashboard is going to be your best idea on information.

[removed]

Next time YOU are about to spout off about something, perhaps think about reading the f'ing page being linked to?

"The issue with connectivity between the GCP us-east1, us-east4, and us-central1 regions to other Cloud Providers has been resolved for all affected projects as of Friday, 2020-03-27 13:37 US/Pacific."

Post reply on HN