Live data from Hacker News

Incident Report: Railway Blocked by Google Cloud [resolved]

status.railway.com

181–190 of 381 posts

Re: Incident Report: Railway Blocked by Google Cloud [resolved]

#181

It has been 0 days since GCP has taken down a startup (again). You see this at least once a year. Never heard of this from AWS or Azure. In all seriousness, this is why we don't use them. They have the most ergonomic cloud of the big three, then absolutely murder it by having this kind of reputation.

On the other hand i can’t remember when there was a serious outage on GCP, unlike AWS/Azure who seem to go down catastrophically a couple of times per year.

>On the other hand i can’t remember when there was a serious outage on GCP

They had a really bad global outage a year ago. At least with AWS outages are contained to a single region.

Re: Incident Report: Railway Blocked by Google Cloud [resolved]

#182
May 2024 UniSuper incident: https://cloud.google.com/blog/products/infrastructure/detail...

https://www.unisuper.com.au/about-us/media-centre/2024/a-joi...

A joint statement from UniSuper CEO Peter Chun and Google Cloud CEO Thomas Kurian

8 May 2024

UniSuper and Google Cloud understand the disruption to services experienced by members has been extremely frustrating and disappointing. We extend our sincere apologies to all members.

While supporting UniSuper to bring its systems back online, Google Cloud has been conducting a root cause analysis.

Thomas Kurian has confirmed that the disruption arose from an unprecedented sequence of events, where an inadvertent misconfiguration during provisioning of UniSuper’s Private Cloud services ultimately resulted in the deletion of UniSuper’s Private Cloud subscription.

This is described as an isolated, “one-of-a-kind occurrence” that has never before occurred with any Google Cloud client globally. This should not have happened. Google Cloud has identified the sequence of events and taken measures to ensure it does not happen again.

Why did the outage last so long?

UniSuper had duplication across two geographies as protection against outages and data loss. However, the deletion of the Private Cloud subscription triggered deletion across both geographies.

Restoring the Private Cloud required significant coordination and effort between UniSuper and Google Cloud, including recovery of hundreds of virtual machines, databases, and applications.

Re: Incident Report: Railway Blocked by Google Cloud [resolved]

#183
post #161
post #131

Earlier quoted context omitted.

Nope, Railway was the company who was hosting PocketOS, which is the company that blamed Cursor for deleting their prod database. Railway is only involved insofar as their API allowed an instant delete of the prod database.

Railway deserves a lot of blame here. Deleting backups along with the database is a lot like not having backups. Moronic design choice.

Why does Railway deserve any blame here at all? It was an MCP with elevated infra access, that the user willingly connected through Cursor, which allowed an LLM Agent to manage infra on Railway. The user would first have gone through oAuth confirming the access level scope (I would have rejected the moment it indicates to me that it can delete critical infra and backups...). So obviously it has access to all commands the user would also have access to. From my perspective the blame is entirely on the user, and partly on Cursor for not enforcing HITL correctly across their agents.

Re: Incident Report: Railway Blocked by Google Cloud [resolved]

#184
post #155

Earlier quoted context omitted.

On the other hand i can’t remember when there was a serious outage on GCP, unlike AWS/Azure who seem to go down catastrophically a couple of times per year.

I've been in AWS for almost twenty years at this point. It's been a long time since I've seen a global outage of the data plane on anything. The control plane, especially the US-east-1 services? Yes - but if you're off of east-1, your outages are measured in missile strikes, not botched deployments.

Didn't the latest outage affect people not on us-east-1 because internal aws services depend on us-east-1?

Re: Incident Report: Railway Blocked by Google Cloud [resolved]

#186

Does anyone know how this even happens inside the walls of google? Is it an automated process? How is such a (presumably) high revenue account just magically blocked without human intervention? I'm quite perplexed.

There would have been efforts to contact them, but it would have been via their contact method, aka the email they set it up with. Common ways this happens? They are using a credit card to run their business with no backup payment method. Then the company's contact person is on vacation. Sign up for terms. It will get you payment terms!

> via their contact method, aka the email they set it up with

If it's anything like AWS, that may be just one of hundreds of emails they send every day, most of which are just noise.

Re: Incident Report: Railway Blocked by Google Cloud [resolved]

#187

How the heck do these things happen, especially with companies with huge monthly spend? At my last job we had some suspicious workloads running on AWS and our TAM reached out to us before taking any action. Who wants to bet this was some AI automation gone wrong and because GCP seems to be allergic to actually contacting a human to get a response, this just sits in some support queue that outsourced workers look at a…

It's Google. They let you use their services, but the moment you don't fit the norm, they suspend you.

What does blocked mean? Is there a different post that I am missing? There is shared infrastructure in GCP for networking (ex-googler here) and if only railway is affected, then it is not clear if it is only GCP or if there is something from Railway's perspective that needs to be addressed.

Re: Incident Report: Railway Blocked by Google Cloud [resolved]

#188

It has been 0 days since GCP has taken down a startup (again). You see this at least once a year. Never heard of this from AWS or Azure. In all seriousness, this is why we don't use them. They have the most ergonomic cloud of the big three, then absolutely murder it by having this kind of reputation.

https://en.wikipedia.org/wiki/Timeline_of_Amazon_Web_Service...

Azure nerfed the front door of all Azure and O365 services last year.

All of these companies are great at what they did, and occasionally fuck up.

Re: Incident Report: Railway Blocked by Google Cloud [resolved]

#189

Earlier quoted context omitted.

On the other hand i can’t remember when there was a serious outage on GCP, unlike AWS/Azure who seem to go down catastrophically a couple of times per year.

Perhaps you don't notice GCP outages because so few companies rely on them?

GCP has a lot of customers. But you wouldn't know the companies that do, unless you worked there and wanted to leak it, or it publicly comes out. Eg it's been publicly acknowledged that Apple uses GCP for iCloud, https://www.cnbc.com/amp/2018/02/26/apple-confirms-it-uses-g... , and Home Depot is another that's used as a case study, https://cloud.google.com/customers/the-home-depot but most customers don't want to make a big deal about being on GCP as it's none of our business who's hosting them.
Post reply on HN