Live data from Hacker News

GCP Outage

status.cloud.google.com

331–340 of 539 posts

Re: GCP Outage

#331

Earlier quoted context omitted.

Doesn't cloudflare have its own infrastructure, it's wild to me that both these things are down presumably together with this size of a blast radius.

You'd think so wouldn't you? DownDetector also reports azure and oracle cloud, I can't see then also being dependant on GCP... I guess down detector isn't a full source of truth though. https://ocistatus.oraclecloud.com/#/ https://azure.status.microsoft/en-gb/status Both green

Down detector has a problem when whole clouds go down: unexpected dependencies. You see an app on a non-problematic cloud is having trouble, and report it to Down Detector but that cloud is actually fine- their actual stuff is running fine. What is really happening is that the app you are using has a dependency on a different SaaS provider who runs on the problematic cloud, and that is killing them.

It's often things like "we got backpressure like we're supposed to, so we gave the end user an error because the processing queue had built up above threshold, but it was because waiting for the timeout from SaaS X slowed down the processing so much that the queue built up." (Have the scars from this more than once.)

Re: GCP Outage

#334

Earlier quoted context omitted.

Why can't companies be honest with being down. It helps us all out so we don't spend an hour internalizing. We are truly in gods hands. $ prod Fetching cluster endpoint and auth data. ERROR: (gcloud.container.clusters.get-credentials) ResponseError: code=503, message=Visibility check was unavailable. Please retry the request and contact support if the problem persists

Because they have unrealistic targets so they make up fake uptime numbers. 99.999% would mean not even having an hour of downtime in 10 years. I remember reddit being down for like a whole day or so and they claimed 99.5% in that month.

Ma Bell hit that decently often.

Re: GCP Outage

#335

Does anyone know of a good dashboard to check for such BGP routing anomalies as (apparently) this one? I am currently digging around https://radar.cloudflare.com/routing but it doesn't show which routes were actually leaked. I would love if anyone has any good tool recommendations!

I am a newb at this too, but is it "normal" for the "Announced IP Address Space" section to have that large jump from addresses like that?

go https://status.gcp.databricks.com/

Re: GCP Outage

#338
Borg and K8s were fighting for resources, so Gemini decided to take out DNS. Now a sysAdmin has to step in.

* just trying to add a little humour. pretty stressfull outage. grarr!!

Re: GCP Outage

#339
post #310

The status page is green, but there are outages reported: https://downdetector.com/status/google-cloud/

Here's the incident: https://status.cloud.google.com/incidents/ow5i3PPK96RduMcb1S...

It was nearly an hour into our company's internal incident channel on this for GCP to finally declare that yes, in fact, things on fire.

… I get that PR-types probably want to massage the message, but going radio dark is not good PR.

Post reply on HN