Live data from Hacker News

Incident Report: Railway Blocked by Google Cloud [resolved]

status.railway.com

271–280 of 381 posts

Re: Incident Report: Railway Blocked by Google Cloud [resolved]

#271
post #197

May 2024 UniSuper incident: https://cloud.google.com/blog/products/infrastructure/detail... https://www.unisuper.com.au/about-us/media-centre/2024/a-joi... A joint statement from UniSuper CEO Peter Chun and Google Cloud CEO Thomas Kurian 8 May 2024 UniSuper and Google Cloud understand the disruption to services experienced by members has been extremely frustrating and disappointing. We extend our sincere apologies to…

The instant cascading worldwide deletion upon closing or deleting a subscription sounds like a recipe for disaster. Why not mark it for deletion and delete say... a day or a week later?

> The instant cascading worldwide deletion upon closing or deleting a subscription sounds like a recipe for disaster.

I don't agree. What do you expect to happen when you explicitly delete your user account? Do you expect your systems to remain in operation for a week? That itself would be a major risk and liability, as your whole infrastructure would still be up even though you cut your access to it.

Also, isn't your whole infrastructure expected to be automatically deployed with IaC? The notable exception is data, which is already soft deleted and recoverable through customer support.

All in all, where do you expect the customer's responsibility to end and the cloud provider's to start? The shared responsibility model is covered by any intro course in no uncertain terms.

Re: Incident Report: Railway Blocked by Google Cloud [resolved]

#272

Earlier quoted context omitted.

upvoted & favourited because you taught me a really interesting fact which I feel makes up for an amazing discussion (regarding icloud using GCP). also, I can't help but imagine if instead of render, it was Apple's account which could've been auto-banned (Render is almost a billion dollar company or series-B, I am not sure) I haven't read the articles and I admit that but can you please elaborate to me on why Apple u…

> So in some sense if Apple is using gcp's for icloud then aren't they just reselling google storage themselves and google can always beat them in pricing while also wanting to chew away at the percentage of iphones themselves too? Apple uses Samsung displays and Sony camera sensors, iirc, both of which are flagship Android phone makers. That doesn't really seem to be a concern in their procurement thinking. iCloud a…

Let's also not ignore enterprise realities: in your example, Samsung Displays is likely giving a great price to Apple for displays based on long-term commitment of large quantities: it allows them to optimize production and possibly give a better price than maybe Samsung Mobile for smaller-runs of phones.

Each division also cross-charges, so Samsung Mobile would be paying Samsung Displays for the screens, and possibly at a small, guaranteed and non-negotiable margin.

Without a global strategy not to do so, divisions within an enterprise optimize for their own bottom line and have internal discussions on build-vs-buy even if they have an internal factory.

Re: Incident Report: Railway Blocked by Google Cloud [resolved]

#273

Earlier quoted context omitted.

It’s a good question. That said unless there are compliance or fallback concerns i would prefer a service that burns my data on departure.

No, that's the naive view Because in case of a compromise/unauthorized access that's exactly what you don't want to happen

> No, that's the naive view

No, not really. That's pretty basic stuff. You would do well in reading up on the shared responsibility model. Customers are responsible for setting up their own infrastructure, and platform/service providers are only responsible for the services they manage. Even then, stuff like persisted data is still recoverable by design.

But you are absolutely responsible for the service you put together. This is a basic principle for around two decades. Infrastructure as code tools are pervasive and ubiquitous for over a decade.

Re: Incident Report: Railway Blocked by Google Cloud [resolved]

#274

Everyone is eager to point a finger at Google, but I've been a user of Railway for a while now, and I've seen enough nonsense to want to hear what GCP has to say about this before drawing any conclusions. Let's just say Railway has had problems like this before, and the way their team handles them does not inspire any confidence. Regardless of how it happened, for me, this is the straw that broke the camel's back.

another ditto from me, albeit anecdotal again. Railway dev teams play fast and loose with sprinkles of vibe coding everywhere on top. There's 'oops yea bear with us we are still a startup' and then there's railway.

Re: Incident Report: Railway Blocked by Google Cloud [resolved]

#275
post #194

Earlier quoted context omitted.

During my 5 years of my startup, we had only 1 outage due to AWS because we picked us-west-2 as the primary reason. If anyone starting a company and picks us-east-1 as the primary reason, they should be fired. There's absolutely no reason to be in that region.

Why do people want to be in that region? Is it the default or something? I know some workloads help to be colocated but all these places are connected by fiber and every cloud has a worldwide CDN it seems.

At some point it used to be significantly cheaper than any other AWS region in the world. Not sure if that's still true.

Re: Incident Report: Railway Blocked by Google Cloud [resolved]

#276

Everyone is eager to point a finger at Google, but I've been a user of Railway for a while now, and I've seen enough nonsense to want to hear what GCP has to say about this before drawing any conclusions. Let's just say Railway has had problems like this before, and the way their team handles them does not inspire any confidence. Regardless of how it happened, for me, this is the straw that broke the camel's back.

> Let's just say Railway has had problems like this before, and the way their team handles them does not inspire any confidence. This. It's very odd that in other threads we see a bunch of accounts heavily invested in criticizing a cloud provider, but what's conspicuously absent from this wave of indignation is any curiosity in the root cause, or even any interest in exploring what it might have been. Quite odd.

Agreed, I'm very curious as to how this could happen.

But TheRegister did reach out to Google and they have not replied yet: https://www.theregister.com/off-prem/2026/05/20/google-cloud...

Re: Incident Report: Railway Blocked by Google Cloud [resolved]

#277

Earlier quoted context omitted.

On the other hand i can’t remember when there was a serious outage on GCP, unlike AWS/Azure who seem to go down catastrophically a couple of times per year.

Perhaps you don't notice GCP outages because so few companies rely on them?

> Perhaps you don't notice GCP outages because so few companies rely on them?

GCP is the world's third largest cloud provider, and has around half of AWS' market share. Claiming no one uses it reads like Yogi Berra's "no one goes there anymore, it's too crowded".

Re: Incident Report: Railway Blocked by Google Cloud [resolved]

#278
post #194

Earlier quoted context omitted.

During my 5 years of my startup, we had only 1 outage due to AWS because we picked us-west-2 as the primary reason. If anyone starting a company and picks us-east-1 as the primary reason, they should be fired. There's absolutely no reason to be in that region.

Why do people want to be in that region? Is it the default or something? I know some workloads help to be colocated but all these places are connected by fiber and every cloud has a worldwide CDN it seems.

> Why do people want to be in that region? Is it the default or something?

It's one of the oldest and largest regions. It hosts the most services, both low-level platform stuff and higher level managed services (which run on the low-level platform stuff), so services tend to be more performant.

Geographic location is also good.

Also, due to scale their pricing ends up being cheaper.

Let's say that it's the region people use by default, unless they have a compelling reason to have a presence in any other particular region.

Re: Incident Report: Railway Blocked by Google Cloud [resolved]

#279

Lesson learned: don't rely on a single hyperscaler, even (or especially) as a startup.

I just... I don't really understand why startups even use AWS, GC, or any other cloud hosted software? Hetzner, etc. Are all extremely cheap, and honestly scale so well... Code nowadays is cheaper for configs, and having full control over your compute is... liberating.

Perhaps Railway does a bit more than what you think, they have some great functionality (I'm not affiliated with them). Check out [Features | Railway](https://railway.com/features) "PR Environments", they are incredible for the QA process

Re: Incident Report: Railway Blocked by Google Cloud [resolved]

#280

Earlier quoted context omitted.

On the other hand i can’t remember when there was a serious outage on GCP, unlike AWS/Azure who seem to go down catastrophically a couple of times per year.

Perhaps you don't notice GCP outages because so few companies rely on them?

There is a mobile game I know of that had an outage as a result of a GCP service outage. That is the only time I've noticed GCP outages.

With that said, I would not say few companies rely on GCP. Search for "GCP" in this month's HN hiring thread. There are 23 hits, more than Azure's 21. AWS has 90 hits, which I guess shows its sheer dominance in the startup space. But these figures more or less agree with my intuition of the major clouds being AWS/GCP/Azure.

Post reply on HN