Live data from Hacker News

Fly.io Postgres cluster down for 3 days, no word from them about it

webcache.googleusercontent.com

311–320 of 493 posts

Re: Fly.io Postgres cluster down for 3 days, no word from them about it

#311
post #80

Y'all, this is going to be deeply unsatisfying, but it's what I can report personally: I have no earthly clue why this thread on our community site is unlisted. We're looking at the admin UI for it right now, and there's like, a little lock next to do the story, but the "unlist story" option is still there for us to click. The best I can say is: I'm reasonably sure there wasn't some top-down edict to hide this thread…

Honest advice, probably to Kurt rather than you, is you need better processes, accountability and (probably) communication in your company. The tone of your reply (and other communications from fly.io) is reflective of the lack of those things given the public sentiment regarding fly.io. At 60+ employees and so many issues that tone goes from humanly endearing to indicative of a non-scaling business. Other replies indicate you don't want the things (process, oversight, etc.) that a growing B2B business needs to really succeed which is not a good sign. Sure there's a cost to that corporate-ness and you want to minimize that cost but it's also a necessary evil for the business you're in at the scale you're at.

If something breaks once it's an accident, if it breaks twice it's bad luck but if it breaks down three times it's broken processes. Based on the comment here things break at fly.io a lot more often than three times.

Re: Fly.io Postgres cluster down for 3 days, no word from them about it

#312

Earlier quoted context omitted.

Ok as long as we’re getting conspiratorial, something similar I observed has bugged me. About a year ago fly awarded a few people in the forums, I think it was 3, the “aeronaut” badge. Basically just pointless bling for a “routinely very helpful” person or somesuch. Still, I can imagine it was cool to get it. No, it wasn’t me. One person I saw with it absolutely deserved it: this person is, to this day, always hoppin…

Get out of here with this nonsense. We tell people when we’re a bad option all the time. Do you really think we have a desire (or time) to punish somebody for doing the same? Also, here’s the long forgotten badge, still with 3 people… https://community.fly.io/badges/107/aeronaut

[flagged]

Re: Fly.io Postgres cluster down for 3 days, no word from them about it

#313

Earlier quoted context omitted.

I moved from DO to Hetzner ( cheaper), I am happy about it.

Does anyone know how Hetzner pricing is half of DO yet is profitable, while DO is loss making with 6% operating margin?

Overstaffed, overinflated and inefficient Silicon Valley startup vs. organically-grown, well-adjusted, efficient German company.

Re: Fly.io Postgres cluster down for 3 days, no word from them about it

#314

There is now a response to the support thread from Fly[1]: > Hi Folks, > Just wanted to provide some more details on what happened here, both with the thread and the host issue. > The radio silence in this thread wasn’t intentional, and I’m sorry if it seemed that way. While we check the forum regularly, sometimes topics get missed. Unfortunately this thread one slipped by us until today, when someone saw it and flag…

total shot in the dark, but, was it a transaction id wrap around?

Re: Fly.io Postgres cluster down for 3 days, no word from them about it

#315
post #152

Earlier quoted context omitted.

> We strongly recommend running multiple instances to mitigate the impact of single-host failures like this. Make it impossible not to do so, and make it frictionless then.

That would presumably cost more money which is not a trade off every user would want to make.

You cannot make every user happy, and its generally better to not have a user than to have an unhappy user.

Re: Fly.io Postgres cluster down for 3 days, no word from them about it

#316
Most people in this thread are either in North America or Europe, so options for managed services like Fly exist, and they are plenty. But for people in South America, what options are there for a Heroku-like service? I don't want users shooting off requests halfway across the globe and back when there are many datacenters a couple of miles from our users. I just don't have the time and resources to manage VMs and scaling issues. I need a zero friction "./serviceX deploy" experience.

Fly seems unreliable, but they offer a deploy region close to me. Does anyone have know of any alternatives?

Re: Fly.io Postgres cluster down for 3 days, no word from them about it

#317
post #257

Earlier quoted context omitted.

If you're deploying an Elixir/Phoenix app, then Gigalixir has worked really well for me. It's expensive, but then so is Heroku.

What’s their reliability been like? Am I right in thinking the platform got bought a little while ago, and it’s being run by a relatively small outfit?

I've been using them for the last ~10 months or so to run http://PhoenixOnRails.com. Gigalixir have been 100% reliable for me so far, but it's a low-traffic app - I can't tell you what it's like to run a big app on them at scale.

I don't know who owns them but I do get the impression it's a small team. Hasn't been an issue for me so far. Their customer service has been very helpful and responsive on the rare occasions I've needed to contact them

Re: Fly.io Postgres cluster down for 3 days, no word from them about it

#318
post #229

Earlier quoted context omitted.

Does anyone know how Hetzner pricing is half of DO yet is profitable, while DO is loss making with 6% operating margin?

They run their own data centres and have for a while. There is a pretty big industry for that sort of thing as an alternative to “the cloud” here in Europe. We used to use nianet to house our hardware in Denmark. Basically these companies does hardware renting and they also do hardware renting with more steps which is where you rent rack space but own the hardware. They provide the place for the hardware and they als…

Hetzner also do some crazy-cool stuff, especially around the 7950X3D, cooling, AM5 etc. (https://www.youtube.com/watch?v=V2P8mjWRqpk). They also do some amazing stuff with ARM (their cloud offering is really solid for this).

Re: Fly.io Postgres cluster down for 3 days, no word from them about it

#319
post #296

There is now a response to the support thread from Fly[1]: > Hi Folks, > Just wanted to provide some more details on what happened here, both with the thread and the host issue. > The radio silence in this thread wasn’t intentional, and I’m sorry if it seemed that way. While we check the forum regularly, sometimes topics get missed. Unfortunately this thread one slipped by us until today, when someone saw it and flag…

I was confused why support for platform failure relies on a forum where employees may or may not check. After checking docs[1], apparently you have to be on a paid plan (at least $29/mo) to access email support, so you may not have it even you’re paying for resources. I won’t be using it for side projects where I’m okay with paying $5-10/mo but don’t want to have three day outages. [1] https://fly.io/docs/about/suppo…

Forewarning: I am not being critical of fly.io nor their free support whatsoever when I say this.

From a technical perspective, could they have "been better" from a technical perspective? I see their name a lot on HN so I know they are doing really cool + advanced things and this is probably some super small edge case that slipped through the cracks.

Could they have added some message / do we as the HN community feel they needed to be like "we're gonna add some extra logging/monitoring going forward so it won't happen again"?

By all means, they probably don't owe anybody in terms of stability + uptime guarantees when it comes to a free tier. Sh*t happens.

Re: Fly.io Postgres cluster down for 3 days, no word from them about it

#320
post #169
post #132

Not a sarcastic or rhetorical question - how come the three big A clouds or even smaller ones (Hetzner,my favorite) are mostly so stable (give or take some outages) and anyone knows their internal engineering, architecture and practices to keep systems that much stable?

There isn't really secret sauce to it in 2023. The techniques, processes, and etc have pretty much been documented over the past 20 years. But if you are wondering how AWS manages to be so good at it at such scale? Hosting infrastructure is incredibly complicated and AWS employs something like 100k people. Seemingly small AWS services employ more engineers than Fly.io. That being said my take is that what's happening…

I think the core issue is that they venomously don't want to act like a corporation. Which is great for early marketing and adoption but there's a reason successful B2B corporations act like they do. It's less fun and it's less endearing but it also annoys customers significantly less. I mean, the CEO has "Interim Food Taster" as his title on LinkedIn.
Post reply on HN