Live data from Hacker News

Fly.io Postgres cluster down for 3 days, no word from them about it

webcache.googleusercontent.com

361–370 of 493 posts

Re: Fly.io Postgres cluster down for 3 days, no word from them about it

#361
post #18

Earlier quoted context omitted.

For what it’s worth, I left Fly because of this crap. At first my Fly machine web app had intermittent connection issues to a new production PG machine. Then my PG machine died. Hard. I lost all data. A restart didn’t work - it could not recover. I restored an older backup over at RDS and couldn’t be happier I left.

I left digitalocean for fly because some of their tooling was excellent. I was pretty excited. I’m back on digitalocean now. I’m not unhappy about it, they’re very solid. I don’t love some things about their services, but overall I’d highly recommend them to other developers. I gave up on fly because I’d spontaneously be unable to automate deployments due to limited resources. Or I’d have previously happy deployments…

I'm a fan of Linode as well.

I want to like Fly, but the reliability is one of those were I feel like every time I investigate moving workloads over I'm disappointed by these stories over and over again.

Re: Fly.io Postgres cluster down for 3 days, no word from them about it

#362
post #296

Earlier quoted context omitted.

I was confused why support for platform failure relies on a forum where employees may or may not check. After checking docs[1], apparently you have to be on a paid plan (at least $29/mo) to access email support, so you may not have it even you’re paying for resources. I won’t be using it for side projects where I’m okay with paying $5-10/mo but don’t want to have three day outages. [1] https://fly.io/docs/about/suppo…

Forewarning: I am not being critical of fly.io nor their free support whatsoever when I say this. From a technical perspective, could they have "been better" from a technical perspective? I see their name a lot on HN so I know they are doing really cool + advanced things and this is probably some super small edge case that slipped through the cracks. Could they have added some message / do we as the HN community feel…

FWIW: I am on the bottom tier of the paid plans ($29/mo) so I could get access to the email support, and even with that their response time is still not great.

I have an ongoing issue with one of my PG clusters where one of the nodes was failing and all my attempts at fixing it are failing (mainly cloning one of the other machines to bring the cluster numbers back to normal).

I emailed my account’s support email mid Friday morning last week and did not hear back until this past Monday night.

Sucks, because like a lot of others in this thread I like what Fly is trying to do and am rooting for them, but IMO they should use a significant chunk of that funding they just received on hiring a ton of SREs and front line customer support.

EDIT: I should add, the past times I have emailed them the response time was good. It's just this most recent time was so egregious (3 days!) to get even that initial response that I bring it up.

Re: Fly.io Postgres cluster down for 3 days, no word from them about it

#363
post #18

There is now a response to the support thread from Fly[1]: > Hi Folks, > Just wanted to provide some more details on what happened here, both with the thread and the host issue. > The radio silence in this thread wasn’t intentional, and I’m sorry if it seemed that way. While we check the forum regularly, sometimes topics get missed. Unfortunately this thread one slipped by us until today, when someone saw it and flag…

For what it’s worth, I left Fly because of this crap. At first my Fly machine web app had intermittent connection issues to a new production PG machine. Then my PG machine died. Hard. I lost all data. A restart didn’t work - it could not recover. I restored an older backup over at RDS and couldn’t be happier I left.

[deleted]

Re: Fly.io Postgres cluster down for 3 days, no word from them about it

#364
Why is this company always on HN frontpage - ironically for their bad services? Normally, poor service from a provider isn't grounds for such attention - but seems like Fly.io has not done anything great.

They still continue to get love from the developer community who "wants them to succeed". I'm puzzled as to why? Because of some blog posts?

Re: Fly.io Postgres cluster down for 3 days, no word from them about it

#366

How many days work is it to build a deployment of an Elixir app with Pulumi, Github Actions and AWS? As someone not incredibly experienced with devops, I always wonder what is best with databases? Should they be provisioned in Pulumi or do I just manually create them in RDS? Secrets Manager seems like a bit of a pain point as does IAM which I think I just about understand until I get lost! Giving everything access to…

It depends how familiar you are. I could probably knock that out in a day with the CDK.

Re: Fly.io Postgres cluster down for 3 days, no word from them about it

#367

Earlier quoted context omitted.

Was the sarcasm an attempt to discredit criticism of Fly's operational processes by pointing out that another company also has issues in how they handle outage notifications?

You could have just answered my previous question with "No, I am not familiar with sarcasm". Because you clearly don't understand sarcasm, I'll be blunt: No, I'm not trying to discredit any criticism of this provider. I agree with the comment I replied to, that this kind of failure mode is fucking ridiculous. My response thus is not an attempt to normalise this, but to highlight the elephant in the room, which is tha…

Thanks for clarifying. I understand now you wanted to call attention to the fact that another famous organization in the same space as Fly.io also has such bad practices. Thanks for the data point.

Re: Fly.io Postgres cluster down for 3 days, no word from them about it

#368
post #99

Earlier quoted context omitted.

(a) Not even close to the first negative HN thread about us. (b) We definitely didn't make the thread private in response to HN. (c) It should be public again.

what's up with the status page?

There's a global status page, and then there's a local update for people with instances on an affected host --- past some threshold of hosts, the probability of having an issue on some random host gets pretty high just because math. The local status thing happened for people with instances on that machine.

Ordinarily, a single-host incident takes a couple minutes to resolve, and, ordinarily, when it's resolved, everything that was running on the host pops right back up. This single-host outage wasn't ordinary. Somehow, a containerd boltdb got corrupted, and it took something like 12 hours for a member of our team (themselves a containerd maintainer) to do some kind of unholy surgery on that database to bring the machine back online.

The runbook we have for handling and communicating single-host outages wasn't tuned for this kind of extended outage. It will be now. Probably we'll just paint the global status page when a single-host outage crosses some kind of time threshold.

Re: Fly.io Postgres cluster down for 3 days, no word from them about it

#369
post #80

Y'all, this is going to be deeply unsatisfying, but it's what I can report personally: I have no earthly clue why this thread on our community site is unlisted. We're looking at the admin UI for it right now, and there's like, a little lock next to do the story, but the "unlist story" option is still there for us to click. The best I can say is: I'm reasonably sure there wasn't some top-down edict to hide this thread…

Honest advice, probably to Kurt rather than you, is you need better processes, accountability and (probably) communication in your company. The tone of your reply (and other communications from fly.io) is reflective of the lack of those things given the public sentiment regarding fly.io. At 60+ employees and so many issues that tone goes from humanly endearing to indicative of a non-scaling business. Other replies in…

I'm just a person on Hacker News that happens to be at Fly.io; as I've said before, it's probably reasonable to think of me as an HN person first, and a Fly.io person second. My tone is my tone, and has been for the many years I've participated in this community. I got back from an evening out, saw that we were on the front page, poked around a little to find out what the hell was going on, and did my best to add some context. That's all.

If you're reading my comments on HN as some kind of official response from the company, you've misconstrued them.

Re: Fly.io Postgres cluster down for 3 days, no word from them about it

#370

Earlier quoted context omitted.

I left digitalocean for fly because some of their tooling was excellent. I was pretty excited. I’m back on digitalocean now. I’m not unhappy about it, they’re very solid. I don’t love some things about their services, but overall I’d highly recommend them to other developers. I gave up on fly because I’d spontaneously be unable to automate deployments due to limited resources. Or I’d have previously happy deployments…

I adore DO. They’re seriously underrated. I love how they’ll just give you a server and say here, have at it. No abstractions, no fancy crap, just get out of my way and let me do my thing.

I went to DO's site due to your comment and I don't see anywhere where I can just get a server. Do you mean a VPS/Droplet? (I'm looking under Products and Solutions.)
Post reply on HN