Live data from Hacker News

Fly.io Postgres cluster down for 3 days, no word from them about it

webcache.googleusercontent.com

1–10 of 493 posts

Re: Fly.io Postgres cluster down for 3 days, no word from them about it

#2
Fly have tried to hush this by making the thread [1] private to anyone not logged in.

One quote from thread:

> This is the second time I’ve had this kind of issue with Fly, where my service just goes down, Fly reports everything healthy, and there’s literally no information and nothing I can really do other than wait and hope it comes back up sometime

Another user:

> We had four machines (app + Postgres for staging and production) running yesterday, and three of the four (including both databases) are still down and can’t be accessed. I can replicate the issues others have mentioned here.

> This is our company’s external API app and so the issue broke all of our integrations.

> Our team ended up setting up a new project in fly to spin up an instance to keep us going which took a couple of hours (backfilling environment variables and configuration etc, not a bad test of our DR ability).

> There is no way I can find to get the data from the db machines. Thank goodness this isn’t our main production db and we were able to reverse engineer what we needed into there.

> Very keen to hear what’s happening with this and why after so many hours there’s no more info or updates.

Another user:

> As an aside, it’s kind of a kick in the teeth to see the status page for our organization reporting no incidents - the same page that lists our apps as under maintenance and inaccessible!

Another user:

> I’m feeling very lucky that none of our paid production apps or databases are affected currently (only our development environment is), but also really surprised that the issue has been ongoing for 17 hours now with no status page update, no notifications (beyond betterstack letting us know it was down) and one note on the app with not much info as to whats going on.

> It really worries me what would happen if it was one of our paid production instances that was affected - the data we’re working with can’t simply be ‘recovered’ later, it’d just get dropped until service resumed or we migrated to another region to get things running again

> Keen to know whats wrong and whats being done about it

Full thread (as at time of HN post; more has been added since): https://pastebin.com/ebmCSZkC

Someone tweeted Fly CEO: https://twitter.com/SouthPawNZ/status/1682181533673857024

[1] https://community.fly.io/t/service-interruption-cant-destroy...

Re: Fly.io Postgres cluster down for 3 days, no word from them about it

#4
I like fly.io a lot and I want them to succeed. They're doing challenging work...things break.

Have to admit it's disappointing to hear about the lack of communication from them, especially when it's something the CEO specifically called out that they wanted to fix in his big reliability post to the community back in March.

https://community.fly.io/t/reliability-its-not-great/11253#s...

Re: Fly.io Postgres cluster down for 3 days, no word from them about it

#7
I think fly.io is pretty incredible but I can't help but feeling they're doomed to follow in heroku's footsteps (unclear if good or bad). They've built some pretty wild stuff and I can't help but wonder if they're overcooking the ocean instead of just solving problems for their users.

Durable and available storage are all they really need to draw me away from big cloud providers but this combined with their answer to S3 being "use S3 or run minio" means I'll never take them seriously.

This is a bad look folks, not sure how you can walk back days of silence and hiding threads. Just open an issue and talk to your users.

Re: Fly.io Postgres cluster down for 3 days, no word from them about it

#9
Wondering if for small/bootstrapped projects there's any alternative people suggest? Fly has a nice UX and accessible prices, but it's unstable at best. I use the big clouds at work, but for personal they are $$$. Also I want to keep devops tending asymptotically to zero.

Re: Fly.io Postgres cluster down for 3 days, no word from them about it

#10

You unfortunately get what you pay for. AWS is more expensive than God, but I'll be damned if you can't have a throat to choke in less than 10 minutes whenever something like this happens.

AWS support replies back to your messages when they feel like it. Their support is just as shady but they have better uptime for sure
Post reply on HN