Live data from Hacker News

GCP Outage

status.cloud.google.com

501–510 of 539 posts

Re: GCP Outage

#501
post #102

Earlier quoted context omitted.

Why can't companies be honest with being down. It helps us all out so we don't spend an hour internalizing. We are truly in gods hands. $ prod Fetching cluster endpoint and auth data. ERROR: (gcloud.container.clusters.get-credentials) ResponseError: code=503, message=Visibility check was unavailable. Please retry the request and contact support if the problem persists

> Why can't companies be honest with being down SLA agreements.

Service level agreements agreements?

Re: GCP Outage

#504
post #471
post #388

Earlier quoted context omitted.

> outage of a 3rd party service that is a key dependency. Good to know that Cloudflare has services seemingly based on GCP with no redundancy.

What's the alternative here? Do you want them to replicate their infrastructure across different cloud providers with automatic fail-over? That sounds -- heck -- I don't know if modern devops is really up to that. It would probably cause more problems than it would solve...

There are roughly 20-25 major IaaS providers in the world that should have close to dependency on each other. I'm almost certain that cloud flare believe that was their posture, and that the action items coming out of this post mortem will be to make sure that this is the case.

Re: GCP Outage

#505
post #388

Earlier quoted context omitted.

> outage of a 3rd party service that is a key dependency. Good to know that Cloudflare has services seemingly based on GCP with no redundancy.

Probably unintentional. "We just read this config from this URL at startup" can easily snowball into "if that URL is unavailable, this service will go down globally, and all running instances will fail to restart when the devops team try to do a pre-emptive rollback"

[deleted]

Re: GCP Outage

#508
post #124

Earlier quoted context omitted.

A program that updates the status page failing does not imply that the status page is manually edited. It is not like you would generate a status page on every request.

How do we know that the program is failing ? How hard is it for the frontend to detect if the last update to the status page was made a while ago and that itself implies there is an error and should be reported ?

We don’t.

But why would the frontend have processing logic when all you need is to serve a static HTML document?

Even if it did, what would you do with that information? Throw up a screen with: Call us for service information at 1-HAHA-JUST-KIDDING

It’s not like it really matters if it’s accurate anyway.

Re: GCP Outage

#509

Earlier quoted context omitted.

GPT is working in agent mode, which kind of confirms that claude is hosted on google and GPT probably on MSFT servers / self hosted.

If you want a stronger confirmation about Claude being hosted on GCP, this is about as authoritative as it gets: https://www.anthropic.com/news/anthropic-partners-with-googl...

That's nearly 2.5 years old, an eternity in this space. It may still be true, but that article is not good evidence.
Post reply on HN