Live data from Hacker News

Incident March 30th, 2026 – Accidental CDN Caching

blog.railway.com

21–30 of 40 posts

Re: Incident March 30th, 2026 – Accidental CDN Caching

#22
I'm kinda shocked (yet not surprised) at how bad railway has been with this:

- Why were they making CDN changes in prod? With their 100M funding recently they could afford a separate env to test CDN changes. Did their engineering team even properly understand surrogate keys to feel confident to roll out a change in prod? I don't think they're beating the AI allegations to figure out CDN configs, a human would not be this confident to test surrogate keys in prod.

- During and post-incident, the comms has been terrible. Initial blog post buried the lede (and didn't even have Incident Report in the title). They only updated this after negative feedback from their customers. I still get the impression they're trying to minimise this, it's pretty dodgy. As other comments mentioned, the post is vague.

- They didn't immediately notify customers about the security incident (people learned from their users). The apparently have emailed affected customers only, many hours after. Some people that were affected that still haven't been emailed, and they seem to be radio silent lately.

- Their founder on twitter keeps using their growth as an excuse for their shoddy engineering, especially lately. Their uptime for what's supposed to be a serious production platform is abysmal, they've clearly prioritised pushing features over reliability https://status.railway.com/ and the issues I've outlined here have little to do with growth, and more to do with company culture.

Honestly, I don't think railway is cut out for real production work (let alone compliance deployments), at least nothing beyond hobby projects.

Their forum is also getting heated, customers have lost revenue, had medical data leaked etc., with no proper followup from the railway team

https://station.railway.com/questions/data-getting-cached-or...

Re: Incident March 30th, 2026 – Accidental CDN Caching

#23
These incidents are a perfect example of how misleading "simple" systems can be.

From the outside, it looks like "just a cache misconfiguration," but in reality, the problem is more insidious because it's distributed across multiple layers: - application logic (authentication limitations) - CDN behavior -> infrastructure - default settings that users rely on (no cache headers because the CDN was disabled)

The hardest part of debugging these cases isn't identifying what happened, but realizing where the model is flawed: everything appears correct locally, the logs don't report any issues, yet users see completely different data.

I've seen similar cases where developers spent hours debugging the application layer before even considering that something upstream was silently changing the behavior.

These are the kind of incidents where the debugging path is anything but linear.

Re: Incident March 30th, 2026 – Accidental CDN Caching

#24
post #19

Earlier quoted context omitted.

> I'm not sure why people are so quick to discount [AI] as a potential source of the issue. Because (per the link above) the CEO said that (1) it was their fault, and (2) it had nothing to do with AI. I understand that on this forum statements like this are inevitably greeted with some amount of skepticism, but right now I'm seeing no particular reason to disbelieve Jake, and the reason that "if they did use AI they'…

Come on man, their CEO is a massive vibe coding proponent and his company spent $300,000 on Claude this month. But yeah, I'm sure Claude had nothing to do with any of it. I bet they don't use it to write any code. https://xcancel.com/JustJake/status/2030063630709096483#m

Both things can be true: they’re doing a lot of vibe coding, and this was a human error that didn’t involve AI.

Re: Incident March 30th, 2026 – Accidental CDN Caching

#25

Earlier quoted context omitted.

Come on man, their CEO is a massive vibe coding proponent and his company spent $300,000 on Claude this month. But yeah, I'm sure Claude had nothing to do with any of it. I bet they don't use it to write any code. https://xcancel.com/JustJake/status/2030063630709096483#m

Both things can be true: they’re doing a lot of vibe coding, and this was a human error that didn’t involve AI.

I have no skin in the game but that is a very charitable perspective.

Re: Incident March 30th, 2026 – Accidental CDN Caching

#26
post #22

I'm kinda shocked (yet not surprised) at how bad railway has been with this: - Why were they making CDN changes in prod? With their 100M funding recently they could afford a separate env to test CDN changes. Did their engineering team even properly understand surrogate keys to feel confident to roll out a change in prod? I don't think they're beating the AI allegations to figure out CDN configs, a human would not be…

Yeah, this was really the nail in the coffin for us. Most services are already moved from Railway, but the rest will follow during this week.

Re: Incident March 30th, 2026 – Accidental CDN Caching

#28
post #22

I'm kinda shocked (yet not surprised) at how bad railway has been with this: - Why were they making CDN changes in prod? With their 100M funding recently they could afford a separate env to test CDN changes. Did their engineering team even properly understand surrogate keys to feel confident to roll out a change in prod? I don't think they're beating the AI allegations to figure out CDN configs, a human would not be…

I was affected and got no communication at all, had to find out from user reports and take immediate action with 0 signal from railway about the issue (even though they were already aware according to the timeline).

I've been trying to defend railway since we built our initial prototype there and I wanted to avoid the cost of migrating to some "serious infra" until proven needed, but they have been making their defense a really hard job (without mentioning that their overall reliability has been really bad the past weeks)

Re: Incident March 30th, 2026 – Accidental CDN Caching

#29
post #19

Earlier quoted context omitted.

Their reply doesn't make much sense, they're supposedly soc2 compliant. How are they compliant but letting a single engineer push out a change like that? I'm sure Claude didn't literally ship the feature itself with no oversight, but I also find it hard to believe that their approach to adopting AI didn't factor in at all. Even just like, the mental overhead of moving faster and adopting AI code with less stringent r…

> I'm not sure why people are so quick to discount [AI] as a potential source of the issue. Because (per the link above) the CEO said that (1) it was their fault, and (2) it had nothing to do with AI. I understand that on this forum statements like this are inevitably greeted with some amount of skepticism, but right now I'm seeing no particular reason to disbelieve Jake, and the reason that "if they did use AI they'…

During this whole incident, Railway have made a wide range of misleading and straight out false claims to cover themselves, so them saying it wasn't AI is pretty much meaningless

Re: Incident March 30th, 2026 – Accidental CDN Caching

#30
post #19

Earlier quoted context omitted.

> I'm not sure why people are so quick to discount [AI] as a potential source of the issue. Because (per the link above) the CEO said that (1) it was their fault, and (2) it had nothing to do with AI. I understand that on this forum statements like this are inevitably greeted with some amount of skepticism, but right now I'm seeing no particular reason to disbelieve Jake, and the reason that "if they did use AI they'…

During this whole incident, Railway have made a wide range of misleading and straight out false claims to cover themselves, so them saying it wasn't AI is pretty much meaningless

Would you mind pointing out these claims? Happy to address them personally
Post reply on HN