Live data from Hacker News

Fly.io Postgres cluster down for 3 days, no word from them about it

webcache.googleusercontent.com

161–170 of 493 posts

Re: Fly.io Postgres cluster down for 3 days, no word from them about it

#161
post #138

What do people get out of using special services like Fly.io instead of standard VMs like the ones you can get from $5/month these days? Can anybody who uses Fly.io explain their rationale? Why do the additional integration with Fly.io, trust and install their special software on your machines and tie your project into their ecosystem? What type of application are you running? How many users are using it?

it's hip, they use hip tech and hired hip folks, so you know it's the place to be ;)

Re: Fly.io Postgres cluster down for 3 days, no word from them about it

#162
post #81

You unfortunately get what you pay for. AWS is more expensive than God, but I'll be damned if you can't have a throat to choke in less than 10 minutes whenever something like this happens.

> a throat to choke yikes

That's a common phrase, not to be taken literally.

It just means one single person (at the vendor) who you can complain to, or raise an issue with.

Re: Fly.io Postgres cluster down for 3 days, no word from them about it

#163
post #68

Earlier quoted context omitted.

Whoa, what? That's a much bigger red flag than the downtime itself.

Ok as long as we’re getting conspiratorial, something similar I observed has bugged me. About a year ago fly awarded a few people in the forums, I think it was 3, the “aeronaut” badge. Basically just pointless bling for a “routinely very helpful” person or somesuch. Still, I can imagine it was cool to get it. No, it wasn’t me. One person I saw with it absolutely deserved it: this person is, to this day, always hoppin…

Get out of here with this nonsense. We tell people when we’re a bad option all the time. Do you really think we have a desire (or time) to punish somebody for doing the same?

Also, here’s the long forgotten badge, still with 3 people… https://community.fly.io/badges/107/aeronaut

Re: Fly.io Postgres cluster down for 3 days, no word from them about it

#164
post #106

Earlier quoted context omitted.

From this my take away is that I could get fired for picking Fly.io for work. Not because there was an outage but because days could pass before getting support. What assurances could you give the community here that the support would be better next time?

This is our public site, for people who don't have support plans with us. It's difficult for me to say more about what happened here and how you might have handled it, because I don't know what happened with this SYD host, because it's 1AM and the people who worked on it are, I assume, asleep. When I know more, I'll do my best to get you a postmortem.

>This is our public site, for people who don't have support plans with us

To be honest, that's enough for me. Sorry I didn't pick up on that.

Re: Fly.io Postgres cluster down for 3 days, no word from them about it

#165
post #151
post #106

Earlier quoted context omitted.

From this my take away is that I could get fired for picking Fly.io for work. Not because there was an outage but because days could pass before getting support. What assurances could you give the community here that the support would be better next time?

Try filing a bug with any of the big three cloud vendors when you're on their free plan. It's really not different, the thing that is going to get you fired is not realizing you're not paying a couple hundred bucks per month for premium service on the infrastructure that is mission critical to your company.

Funny story, when I started my current role I researched our hosting provider. I couldn't find the matching invoices in the accounting system. So I called the vendor, a local company. They'd not set our account up correctly, billing was not enabled. Since then we've been billed. I'm glad we sorted it but it wasn't a good look to start my role by increasing our spending.

Re: Fly.io Postgres cluster down for 3 days, no word from them about it

#167
post #68

Earlier quoted context omitted.

Whoa, what? That's a much bigger red flag than the downtime itself.

Ok as long as we’re getting conspiratorial, something similar I observed has bugged me. About a year ago fly awarded a few people in the forums, I think it was 3, the “aeronaut” badge. Basically just pointless bling for a “routinely very helpful” person or somesuch. Still, I can imagine it was cool to get it. No, it wasn’t me. One person I saw with it absolutely deserved it: this person is, to this day, always hoppin…

Conspiratorial or not that's enough for me to never use it. God forbid someone recommends another platform that handles your clear shortcomings.

Re: Fly.io Postgres cluster down for 3 days, no word from them about it

#168
post #79

Fly works on a new and interesting platform paradigm by writing a lot of their software stack from scratch. Unfortunately, such an approach is unlikely to produce the same stability you might be used to from other places.

I think their proxy could have been written from scratch. Some management, billing, API etc too but under the hood, it's all standard open source stuff like kvm, firecracker and such?

Re: Fly.io Postgres cluster down for 3 days, no word from them about it

#169
post #132

Not a sarcastic or rhetorical question - how come the three big A clouds or even smaller ones (Hetzner,my favorite) are mostly so stable (give or take some outages) and anyone knows their internal engineering, architecture and practices to keep systems that much stable?

There isn't really secret sauce to it in 2023. The techniques, processes, and etc have pretty much been documented over the past 20 years.

But if you are wondering how AWS manages to be so good at it at such scale? Hosting infrastructure is incredibly complicated and AWS employs something like 100k people. Seemingly small AWS services employ more engineers than Fly.io.

That being said my take is that what's happening at Fly.io is a lack of leadership. There are not the right people in the right positions clearly. I've worked infra at companies from 5 people to, well Rackspace, and I'm having a hard time imagining so much time passing with.. Essentially a piece of infra MIA and impacting users.

Re: Fly.io Postgres cluster down for 3 days, no word from them about it

#170
post #81

Earlier quoted context omitted.

> a throat to choke yikes

That's a common phrase, not to be taken literally. It just means one single person (at the vendor) who you can complain to, or raise an issue with.

yea, it's just another one to add to a list of expressions that are unnecessarily aggressive, and for which there are better alternatives
Post reply on HN