Live data from Hacker News

GCP Outage

status.cloud.google.com

511–520 of 539 posts

Re: GCP Outage

#512
post #471
post #388

Earlier quoted context omitted.

> outage of a 3rd party service that is a key dependency. Good to know that Cloudflare has services seemingly based on GCP with no redundancy.

What's the alternative here? Do you want them to replicate their infrastructure across different cloud providers with automatic fail-over? That sounds -- heck -- I don't know if modern devops is really up to that. It would probably cause more problems than it would solve...

> What's the alternative here? Do you want them to replicate their infrastructure

Cloudflare adverises themselves as _the_ redundancy / CDN provider. Don't ask me for an "alternative" but tell them to get their backend infra shit in order.

Re: GCP Outage

#513

Earlier quoted context omitted.

The status page is essentially an admission of guilt. It can require approval from the legal department and a high level official from the company to approve updating it and the verbiage used on the status page.

> It can require approval from the legal department and a high level official from the company to approve updating it and the verbiage used on the status page. Is that true in this case or are you speculating? My company runs a cloud platform. Our strategy is to have outages happen as rarely as possible and to proactively offer rebates based on customer-measured downtime. I don't know why people would trust vendors t…

I don't have any special knowledge about the companies involved in this outage. I do know most (all?) status pages for large companies have to be manually updated and not just anybody can do that. These things impact contracts, so you want to be really sure it is accurate and an actual outage (not just a monitor going off, possibly giving a false positive).

Re: GCP Outage

#514

Earlier quoted context omitted.

What makes you think it’s hard? We have AI generating songs and writing code, but setting up basic health checks is too much?

Yes. “Basic health checks” is not a real thing. I mean that genuinely. > What makes you think it’s hard? Being responsible (or rather, on a team of people responsible) for a status page of a big tech co made me think it’s hard. “Is it down?” Is not a binary question.

[deleted]

Re: GCP Outage

#516
post #486

Earlier quoted context omitted.

I was really surprised. The dependence on another enterprise’s cloud services in-general I think is risky, but pretty much everyone does it these days, but I didn’t expect them to be.

well at some level you can contract deploy private instances of clouds as well.

AWS has Outpost racks that let you run AWS instances and services in your own datacenter managed like the ones running in AWS datacenters. Neat but incredibly expensive.

Re: GCP Outage

#517
post #439

Just our bi-yearly reminder of our over reliance on cloud providers for literally everything. Can't say there's an answer beyond trying to build more independent tech but we know how that goes.

Yet migration to the cloud continues, driven by people arguing that doing it yourself is too complicated and expensive. Let’s see how long until one outage takes down the global economy for multiple days or weeks.

Re: GCP Outage

#519
post #502

Earlier quoted context omitted.

AWS and Azure both had outages.

Is that true? I see no direct report about that. downdetector says so, but it's crowdsourced so it tends to have fake positives.

That's fair, I haven't seen any posts from the companies themselves.

Re: GCP Outage

#520
post #388
post #371

Earlier quoted context omitted.

From the Cloudflare incident: > Cloudflare’s critical Workers KV service went offline due to an outage of a 3rd party service that is a key dependency. As a result, certain Cloudflare products that rely on KV service to store and disseminate information are unavailable [...] Surprising, but not entirely unplausible for a GCP outage to spread to CF.

> outage of a 3rd party service that is a key dependency. Good to know that Cloudflare has services seemingly based on GCP with no redundancy.

After reading about cloudflare infra in post mortems it has always been surprising how immature their stack is. Like they used to run their entire global control plane in a single failure domain.

Im not sure who is running the show there, but the whole thing seems kinda shoddy given cloudflares position as the backbone of a large portion of the internet.

I personally work at a place with less market cap than cloudflare and we were hit by the exact same instances (datacenter power went out) and had almost no downtime, whereas the entire cloudflare api was down for nearly a day.

Post reply on HN