Live data from Hacker News

Fly.io Postgres cluster down for 3 days, no word from them about it

webcache.googleusercontent.com

381–390 of 493 posts

Re: Fly.io Postgres cluster down for 3 days, no word from them about it

#381
post #151
post #106

Earlier quoted context omitted.

From this my take away is that I could get fired for picking Fly.io for work. Not because there was an outage but because days could pass before getting support. What assurances could you give the community here that the support would be better next time?

Try filing a bug with any of the big three cloud vendors when you're on their free plan. It's really not different, the thing that is going to get you fired is not realizing you're not paying a couple hundred bucks per month for premium service on the infrastructure that is mission critical to your company.

> Try filing a bug with any of the big three cloud vendors when you're on their free plan.

A host being down for 3 days isn’t a bug. And you can contact AWS support, even on the free plan, and get a reply. Try it yourself. The great thing about AWS and the other cloud providers? If a host has issues they email all customers with workloads on it so you don’t need to refresh or check a forum.

I understand fly is a community darling. They’re unreliable, with poor support currently. Maybe the dev experience is great and that makes up for it, but pretending like everything else is equally shitty? Not true.

Re: Fly.io Postgres cluster down for 3 days, no word from them about it

#382

Earlier quoted context omitted.

Fly is in my “try later book” from a year or two ago. I remember it was hard to deploy anything due to downtime so gave up. Sad that stuff like this still happens. You shouldn’t need to multi region a postgres yourself - they should have at least 2 data centre redundancy for the region and it just works. Hope they get some magic sauce to become better at this.

> Hope they get some magic sauce to become better at this. When I saw them describe their multiregion SQL replication architecture I thought "what crazy person thought this wouldn't eventually open up a spider's nest of distributed systems errors?"

Our multiregion SQL replication architecture is the standard Postgres multiregion replication architecture. We do single-write-leader, multiple reader replicas, like everybody else does.

Re: Fly.io Postgres cluster down for 3 days, no word from them about it

#383

There's a lot of bullshit in this HN thread, but here's the important takeaway: - it seems their staff were working on the issue before customers noticed it. - once paid support was emailed, it took many hours for them to respond. - it took about 20 hours for an update from them on the downed host. - they weren't updating their users that were affected about the downed host or ways to recover. - the status page was b…

Not a great summary from my perspective. Here's what I got out of it: - Their free tier support depended on noticing message board activity and they didn't. - Those experiencing outages were seeing the result of deploying in a non-HA configuration. Opinions differ as to whether they were properly aware that they were in that state. - They had an unusually long outage for one particular server. - Those points combined…

I'm surprised by your risk tolerance. If I had any cloud service at this level in my stack go down for three days, I'd start shopping for an alternative. This exceeds the level of acceptability for me for even non-HA requirements. After all, if I can't trust them for this, why would I ever consider giving them my HA business? Just based on napkin math for us, this could've been a potential loss of nearly half a million dollars. Up until this point, I've looked at Fly.io's approach to PR and their business as unconventional but endearing. Now I'm beginning to look at them as unserious. I'm sorry if that sounds harsh. It's the cold truth.

Re: Fly.io Postgres cluster down for 3 days, no word from them about it

#384
post #118

Earlier quoted context omitted.

Is that even possible on Fly?

He may have been talking about Fly themselves. Certainly having only a single machine to serve a wealthy metropolis of 8 million people seems like amateur hour.

Obviously, we have a bunch of machines, both workers and edge servers, in Sydney. The whole Sydney region didn't go down; one worker did.

Re: Fly.io Postgres cluster down for 3 days, no word from them about it

#385

Earlier quoted context omitted.

Not a great summary from my perspective. Here's what I got out of it: - Their free tier support depended on noticing message board activity and they didn't. - Those experiencing outages were seeing the result of deploying in a non-HA configuration. Opinions differ as to whether they were properly aware that they were in that state. - They had an unusually long outage for one particular server. - Those points combined…

Apologies for repeating myself, but: You get to a certain number of servers and the probability on any one day that some server somewhere is going to hiccup and bounce gets pretty high. That's what happened here: a single host in Sydney, one of many, had a problem. When we have an incident with a single host, we update a notification channel for people with instances on that host. They are a tiny sliver of all our us…

Why are your customers exposed to this? This sounds like a tough problem that I'm sympathetic to for you personally, but it sounds like there's no failover or appropriate redundancy in place to rollover to while you work to fix the problem.

edit: I hope this comment doesn't sound accusatory. At the end of the day I want everyone to succeed. I hope there's a silver lining to this in the post-mortem.

Re: Fly.io Postgres cluster down for 3 days, no word from them about it

#386

Earlier quoted context omitted.

Honest advice, probably to Kurt rather than you, is you need better processes, accountability and (probably) communication in your company. The tone of your reply (and other communications from fly.io) is reflective of the lack of those things given the public sentiment regarding fly.io. At 60+ employees and so many issues that tone goes from humanly endearing to indicative of a non-scaling business. Other replies in…

I'm just a person on Hacker News that happens to be at Fly.io; as I've said before, it's probably reasonable to think of me as an HN person first, and a Fly.io person second. My tone is my tone, and has been for the many years I've participated in this community. I got back from an evening out, saw that we were on the front page, poked around a little to find out what the hell was going on, and did my best to add som…

> If you're reading my comments on HN as some kind of official response from the company, you've misconstrued them.

For what it’s worth, this is the reason most companies eventually restrict their employees from making statements about the company; It doesn’t matter if you thought it was clear that is was unofficial, any statement from an employee in a position of power (such as someone with access to the control panel) will be perceived as a communication from the company.

You may have intended it to be a personal remark about your job, but there are a lot of people in this thread looking for any communication they can get about the company.

When you step in to fill that void as a person who appears to have access and power within the company, you are the official communication whether you intend to be or not.

Re: Fly.io Postgres cluster down for 3 days, no word from them about it

#387
post #81

Earlier quoted context omitted.

> a throat to choke yikes

That's a common phrase, not to be taken literally. It just means one single person (at the vendor) who you can complain to, or raise an issue with.

Been in the industry a long time. It’s not a common phrase. It’s weirdly violent. At most “someone to yell at”. A throat to choke? What the fuck.

Re: Fly.io Postgres cluster down for 3 days, no word from them about it

#388

Earlier quoted context omitted.

I adore DO. They’re seriously underrated. I love how they’ll just give you a server and say here, have at it. No abstractions, no fancy crap, just get out of my way and let me do my thing.

I went to DO's site due to your comment and I don't see anywhere where I can just get a server. Do you mean a VPS/Droplet? (I'm looking under Products and Solutions.)

Not GP, but yes -- Droplets are DigitalOcean's "servers" (virtual, but nonetheless).

You boot one up in less than 30 seconds, and get ssh access to it almost immediately. It's very BS-free.

Re: Fly.io Postgres cluster down for 3 days, no word from them about it

#389

There's a lot of bullshit in this HN thread, but here's the important takeaway: - it seems their staff were working on the issue before customers noticed it. - once paid support was emailed, it took many hours for them to respond. - it took about 20 hours for an update from them on the downed host. - they weren't updating their users that were affected about the downed host or ways to recover. - the status page was b…

I've personally had this experience with Fly on a personal project. My project went down but their status pages said everything was up. It's fine since it's personal for fun project but for anything more serious I don't know if I'd be comfortable using them.

Re: Fly.io Postgres cluster down for 3 days, no word from them about it

#390
post #385

Earlier quoted context omitted.

Apologies for repeating myself, but: You get to a certain number of servers and the probability on any one day that some server somewhere is going to hiccup and bounce gets pretty high. That's what happened here: a single host in Sydney, one of many, had a problem. When we have an incident with a single host, we update a notification channel for people with instances on that host. They are a tiny sliver of all our us…

Why are your customers exposed to this? This sounds like a tough problem that I'm sympathetic to for you personally, but it sounds like there's no failover or appropriate redundancy in place to rollover to while you work to fix the problem. edit: I hope this comment doesn't sound accusatory. At the end of the day I want everyone to succeed. I hope there's a silver lining to this in the post-mortem.

The way to not be exposed to this is to run an HA configuration with more than one instance.

If you're running an app on Fly.io without local durable storage, then it's easy to fail over to another server. But durable storage on Fly.io is attached NVMe storage.

By far the most common way people use durable storage on Fly.io is with Postgres databases. If you're doing that on Fly.io, we automatically manage failover at the application layer: you run multiple instances, they configure themselves in a single-writer multi-reader cluster, and if the leader fails, a replica takes over.

We will let you run a single-instance Postgres "cluster", and people definitely do that. The downside to that configuration is, if the host you're on blows up, your availability can take a hit. That's just how the platform works.

Post reply on HN