Earlier quoted context omitted.
I'm not sure on this, will it make any sense - customers who DON'T WANT to be aware of what is required for HA (say lonely devs) choosing such a hosting types. Even if you put educational articles, I'm unsure it will be used. Putting some BANNER IN RED LETTERS into CLI output + link to article may work, though. What do you think?
This is exactly how it currently works: $ fly volumes create mydata Warning! Individual volumes are pinned to individual hosts. You should create two or more volumes per application. You will have downtime if you only create one. Learn more at https://fly.io/docs/reference/volumes/ ? Do you still want to use the volumes feature? (y/N) (and yes, the warning is already even in red letters too)
Fly.io Postgres cluster down for 3 days, no word from them about it
431–440 of 493 posts
Re: Fly.io Postgres cluster down for 3 days, no word from them about it
#432Earlier quoted context omitted.
Honest advice, probably to Kurt rather than you, is you need better processes, accountability and (probably) communication in your company. The tone of your reply (and other communications from fly.io) is reflective of the lack of those things given the public sentiment regarding fly.io. At 60+ employees and so many issues that tone goes from humanly endearing to indicative of a non-scaling business. Other replies in…
I'm just a person on Hacker News that happens to be at Fly.io; as I've said before, it's probably reasonable to think of me as an HN person first, and a Fly.io person second. My tone is my tone, and has been for the many years I've participated in this community. I got back from an evening out, saw that we were on the front page, poked around a little to find out what the hell was going on, and did my best to add som…
Re: Fly.io Postgres cluster down for 3 days, no word from them about it
#433Earlier quoted context omitted.
I'm just a person on Hacker News that happens to be at Fly.io; as I've said before, it's probably reasonable to think of me as an HN person first, and a Fly.io person second. My tone is my tone, and has been for the many years I've participated in this community. I got back from an evening out, saw that we were on the front page, poked around a little to find out what the hell was going on, and did my best to add som…
It seems you took my comment personally but it was about not just your comments but the overall tone of the fly.io communication (see recent blog post regarding funding) and approach to issues (three days of silence on a dead instance). You view processes and guidelines as chains versus as a ladder to help you climb a cliff. If the processes and communication was good then you'd know when you should self-restrict and…
Re: Fly.io Postgres cluster down for 3 days, no word from them about it
#434Earlier quoted context omitted.
I'm not sure on this, will it make any sense - customers who DON'T WANT to be aware of what is required for HA (say lonely devs) choosing such a hosting types. Even if you put educational articles, I'm unsure it will be used. Putting some BANNER IN RED LETTERS into CLI output + link to article may work, though. What do you think?
I agree, articles tend not to get read by those who need them most. A warning from the CLI and a banner on the app management page with a link to a detailed explanation would seem like a good approach. edit: sibling post shows there is such a message on the CLI. The only other thing I can think of is an "Are you sure you want to do this?" prompt, but in the end you can't reach everybody.
Re: Fly.io Postgres cluster down for 3 days, no word from them about it
#435Earlier quoted context omitted.
I'm not sure on this, will it make any sense - customers who DON'T WANT to be aware of what is required for HA (say lonely devs) choosing such a hosting types. Even if you put educational articles, I'm unsure it will be used. Putting some BANNER IN RED LETTERS into CLI output + link to article may work, though. What do you think?
I agree, articles tend not to get read by those who need them most. A warning from the CLI and a banner on the app management page with a link to a detailed explanation would seem like a good approach. edit: sibling post shows there is such a message on the CLI. The only other thing I can think of is an "Are you sure you want to do this?" prompt, but in the end you can't reach everybody.
Re: Fly.io Postgres cluster down for 3 days, no word from them about it
#436Most people in this thread are either in North America or Europe, so options for managed services like Fly exist, and they are plenty. But for people in South America, what options are there for a Heroku-like service? I don't want users shooting off requests halfway across the globe and back when there are many datacenters a couple of miles from our users. I just don't have the time and resources to manage VMs and sc…
Re: Fly.io Postgres cluster down for 3 days, no word from them about it
#437Wondering if for small/bootstrapped projects there's any alternative people suggest? Fly has a nice UX and accessible prices, but it's unstable at best. I use the big clouds at work, but for personal they are $$$. Also I want to keep devops tending asymptotically to zero.
Re: Fly.io Postgres cluster down for 3 days, no word from them about it
#438Earlier quoted context omitted.
"it depends". Dell is fairly good overall, on-site techs are outsourced subcontractors a lot so that can be a mixed bag, pushy sales. Supermicro is good on a budget, not quite mature full fault management or complete SNMP or redfish, they can EOL a new line of gear suddenly.
Have you come across Fujitsu PRIMERGY servers before? https://www.fujitsu.com/global/products/computing/servers/pr... I used to use them a few years ago in a local data centre, and they were pretty good back then. They don't seem to be widely known about though.
Re: Fly.io Postgres cluster down for 3 days, no word from them about it
#439There's a lot of bullshit in this HN thread, but here's the important takeaway: - it seems their staff were working on the issue before customers noticed it. - once paid support was emailed, it took many hours for them to respond. - it took about 20 hours for an update from them on the downed host. - they weren't updating their users that were affected about the downed host or ways to recover. - the status page was b…
>> I get that due to the nature of their plans and architecture, downtime like this is guaranteed and normal. What other cloud providers have downtimes of 20 hours? There must be a lot to call this "guaranteed and normal".
Sadly, I've always felt a good amount of passive aggressiveness in many of the HN threads where fly.io is involved.
Re: Fly.io Postgres cluster down for 3 days, no word from them about it
#440Earlier quoted context omitted.
(Fly.io employee here) To clarify, we communicated this incident to the personalized status page [1] of all affected customers within 30 minutes of this single host going down, and resolved the incident on the status page once it was resolved ~47h later. Here's the timeline (UTC): - 2023-07-17 16:19 - host goes down - 2023-07-17 16:49 - issue posted to personalized status page - 2023-07-19 15:00 - host is fixed - 202…
Ouch? The bad news is that I'd be out of a job if I chose your service in this instance. 47 hours is two full days. For an entire cluster to be down for that long is just unacceptable. Rebuilding a cluster from the last-known-good backup should not take that long, unless there are PBs of data involved; dividing such large data stores into separate clusters/instances seems warranted. Solution archs should steer custom…
> Rebuilding a cluster from the last-known-good backup should not take that long
It's not even clear if that's the right thing to do as a service provider.
Let's say you host a database on some database service, and the entire host is lost. I don't think you want the service provider to restore automatically from the last backup because it makes assumptions about what data loss you're tolerant to. If it just works from the last backup, suddenly you're potentially missing a day of transactions that you thought were there that magically disappears as opposed to knowing they disappeared from a hard break.