Live data from Hacker News

Fly.io Postgres cluster down for 3 days, no word from them about it

webcache.googleusercontent.com

481–490 of 493 posts

Re: Fly.io Postgres cluster down for 3 days, no word from them about it

#481

Why is this company always on HN frontpage - ironically for their bad services? Normally, poor service from a provider isn't grounds for such attention - but seems like Fly.io has not done anything great. They still continue to get love from the developer community who "wants them to succeed". I'm puzzled as to why? Because of some blog posts?

They have really good tech blog posts. Also, they have https://fly.io/dist-sys/

Re: Fly.io Postgres cluster down for 3 days, no word from them about it

#482

Earlier quoted context omitted.

They're not responsible for extreme data recovery, but (almost?) all of the customer data volumes on that server were completely intact. They damn well should be responsible for getting that data back to their customers, whether or not they get the server going again. If you run off a single drive, and the drive dies, any resulting data loss is your fault. But not if something else dies.

I'm absolutely 100% certain that AWS (for example) wouldn't do that for you with the instance types that feature direct attached storage.

Directly attached storage in AWS is a special niche that disappears when you so much as hibernate. And even then they talk about how disk failure loses the data but power failure won't.

This is much closer to EBS breaking. It happens sometimes, but if the data is easily accessible then it shouldn't get tossed.

Re: Fly.io Postgres cluster down for 3 days, no word from them about it

#483

Earlier quoted context omitted.

> I think the assumption of badfaith in this thread from Fly is pretty unprofessional. Customers aren't supposed to show professionalism. Service providers are. I didn't see disrespectful comments here. People here are just poiting this has happened many times and look like a pattern. If you don't fix a communication issue after multiple occurrences, you might not be ill intentioned, but at least careless.

(Note: I edited my comment to make it clear I'm referring to badfaith from commenters, quote above is from pre-edit) I'd argue the expectation goes both ways. I won't link to specific comments, but I think it's pretty clear that some of them cross the line to disrespectful.

I haven't seen a single instance of disrespect -- just justified frustration for the outage and lack of communication. There seem to be a lot of frustrated Fly.io customers and a vocal number of Fly.io fans.

Re: Fly.io Postgres cluster down for 3 days, no word from them about it

#484

Earlier quoted context omitted.

(Note: I edited my comment to make it clear I'm referring to badfaith from commenters, quote above is from pre-edit) I'd argue the expectation goes both ways. I won't link to specific comments, but I think it's pretty clear that some of them cross the line to disrespectful.

I haven't seen a single instance of disrespect -- just justified frustration for the outage and lack of communication. There seem to be a lot of frustrated Fly.io customers and a vocal number of Fly.io fans.

I wonder if there would be anyone defending the service provider if it was AWS or Azure, for instance, instead of Fly.io

Re: Fly.io Postgres cluster down for 3 days, no word from them about it

#485

Earlier quoted context omitted.

(Fly.io employee here) To clarify, we communicated this incident to the personalized status page [1] of all affected customers within 30 minutes of this single host going down, and resolved the incident on the status page once it was resolved ~47h later. Here's the timeline (UTC): - 2023-07-17 16:19 - host goes down - 2023-07-17 16:49 - issue posted to personalized status page - 2023-07-19 15:00 - host is fixed - 202…

Dude. I don't sit at home refreshing status pages. Send me an e-mail. That's how other [useful] providers notify their customers that one of their hosts went down unexpectedly. Linode will send me 6 emails when they need to reboot something. Even Oracle sends me notices about network blips. I believe I've gotten one from AWS, but I also know sometimes their gear gets stuck in a bad state and I didn't get a notificati…

How do you know emails weren't set in addition to the status page changes?

Re: Fly.io Postgres cluster down for 3 days, no word from them about it

#486
post #223

Earlier quoted context omitted.

They don’t.

Where do DO get their servers and data centers from? ... Apparently they run on AWS, I'm surprised

> Apparently they run on AWS, I'm surprised

They don't run on AWS. Not sure what sort of rumors are running :(

> data centers from?

The major players e.g. Equinix, Coresite, etc. Varies per location. Even AWS don't build most of their data centers.

Re: Fly.io Postgres cluster down for 3 days, no word from them about it

#487

Earlier quoted context omitted.

I left digitalocean for fly because some of their tooling was excellent. I was pretty excited. I’m back on digitalocean now. I’m not unhappy about it, they’re very solid. I don’t love some things about their services, but overall I’d highly recommend them to other developers. I gave up on fly because I’d spontaneously be unable to automate deployments due to limited resources. Or I’d have previously happy deployments…

I adore DO. They’re seriously underrated. I love how they’ll just give you a server and say here, have at it. No abstractions, no fancy crap, just get out of my way and let me do my thing.

I find myself going to DO docs on various setup things even when I'm not using said thing on DO (although I'm also a DO customer, and love them for the reasons you've stated).

Re: Fly.io Postgres cluster down for 3 days, no word from them about it

#488
post #80

Y'all, this is going to be deeply unsatisfying, but it's what I can report personally: I have no earthly clue why this thread on our community site is unlisted. We're looking at the admin UI for it right now, and there's like, a little lock next to do the story, but the "unlist story" option is still there for us to click. The best I can say is: I'm reasonably sure there wasn't some top-down edict to hide this thread…

I really like the work that you're doing Thomas, this is the right approach. FWIW, https://fly.io/blog/carving-the-scheduler-out-of-our-orchest... is one of my favourite posts on your blog.

For everyone else reading this, we have been running https://changelog.com on Fly.io since April 2022. This is what our architecture currently looks like: https://github.com/thechangelog/changelog.com/blob/master/IN...

After 15 months & more than 100 million requests served by our Phoenix + PostgreSQL app running on Fly.io, I would be hard pressed to find a reason to complain. - Some deploys failed, and re-running the pipeline fixed it. - Early July 2023, 9k requests from Frankfurt returned 503s. Issue lasted 10 seconds. - While experimenting with machines, after many creations & deletions, one volume could not be deleted. Next day, the volume was gone.

That's about it after 15 months of running production workloads on Fly.io.

We mention about our Fly.io experience often in our Kaizen pod episodes, which we publish every ~2 months: https://changelog.com/topic/kaizen. For anyone curious, this is the episode in which we announced the migration: https://changelog.com/shipit/50. There is a detailed PR which goes with it: https://github.com/thechangelog/changelog.com/pull/407. We've been talking about our migration plan from apps v1 (Nomad) to apps v2 (flyd) recently: https://changelog.com/friends/2#transcript-138

I'm sorry to hear that many of you didn't have the best experience. I know that things will continue improving at Fly.io. My hope is that one day, all these hard times will make for great stories. This gives me hope: https://community.fly.io/t/reliability-its-not-great/11253

Keep improving.

Re: Fly.io Postgres cluster down for 3 days, no word from them about it

#489

Earlier quoted context omitted.

Dude. I don't sit at home refreshing status pages. Send me an e-mail. That's how other [useful] providers notify their customers that one of their hosts went down unexpectedly. Linode will send me 6 emails when they need to reboot something. Even Oracle sends me notices about network blips. I believe I've gotten one from AWS, but I also know sometimes their gear gets stuck in a bad state and I didn't get a notificati…

How do you know emails weren't set in addition to the status page changes?

The whole point of this HN thread is customers weren't getting regular updates. If they had they wouldn't be on a random community forum trying to get support's attention.

Re: Fly.io Postgres cluster down for 3 days, no word from them about it

#490

Earlier quoted context omitted.

Only complaint with Hetzner is they don't have some kind of OAuth setup for machines or scoped API tokens, just read/write. I'd like to use the former for doing Vault authentication from instances, and the latter for writing a dynamic Vault secret provider.

Can’t you use a third party IAM solution for this? Like Okta or keycloak?

zitadel supports service users with rbac. maybe give it a look/try: https://github.com/zitadel/zitadel
Post reply on HN