Live data from Hacker News

Fly.io Postgres cluster down for 3 days, no word from them about it

webcache.googleusercontent.com

421–430 of 493 posts

Re: Fly.io Postgres cluster down for 3 days, no word from them about it

#421

Earlier quoted context omitted.

Honest advice, probably to Kurt rather than you, is you need better processes, accountability and (probably) communication in your company. The tone of your reply (and other communications from fly.io) is reflective of the lack of those things given the public sentiment regarding fly.io. At 60+ employees and so many issues that tone goes from humanly endearing to indicative of a non-scaling business. Other replies in…

I'm just a person on Hacker News that happens to be at Fly.io; as I've said before, it's probably reasonable to think of me as an HN person first, and a Fly.io person second. My tone is my tone, and has been for the many years I've participated in this community. I got back from an evening out, saw that we were on the front page, poked around a little to find out what the hell was going on, and did my best to add som…

TBH I thought you were replying as the CEO of fly.io since 1) I've seen them post here before, 2) I have no idea how big fly.io's staff is and 3) your post didn't otherwise describe who you were. It doesn't look like I was the only one to be confused.

If you had said "thoughts are my own; I just work there" or something I think it would have been more clear.

Re: Fly.io Postgres cluster down for 3 days, no word from them about it

#422
post #242
post #229

Earlier quoted context omitted.

They run their own data centres and have for a while. There is a pretty big industry for that sort of thing as an alternative to “the cloud” here in Europe. We used to use nianet to house our hardware in Denmark. Basically these companies does hardware renting and they also do hardware renting with more steps which is where you rent rack space but own the hardware. They provide the place for the hardware and they als…

commendable to plan a few years ahead, but betting on the state of cloud business 26years from now seems a bit over the top

I think you might misunderstand me. The 2050 is a guesstimate and it's just my opinion on the matter. As far as planning ahead goes, you plan for 5-10 years when you try to figure out where to "iron" your enterprise IT. This is because that's how long your hardware will last if you go the route of renting rack space with your own hardware. I think we tend to plan for 8 years, with some space for "unintended" early failures on things like controllers after 4 years. So while you can contract big-cloud vendors for shorter, I think ours is on 3 year contracts right now, you still sort of do the business case for much longer. Maybe not every 3 years, but at least every 6 years.

You do the same on the other side of the table. Companies like Hetzner knows that EU cloud sollutions are likely to see growth, so it's only natural that they invest in the tech to put themselves in a prime position to jump on the opportunity. Selling a good product while you do so is the way I would do it personally, but you also have EU cloud initiatives backed by VC money going straight for the endgame.

Re: Fly.io Postgres cluster down for 3 days, no word from them about it

#423
post #383

Earlier quoted context omitted.

Not a great summary from my perspective. Here's what I got out of it: - Their free tier support depended on noticing message board activity and they didn't. - Those experiencing outages were seeing the result of deploying in a non-HA configuration. Opinions differ as to whether they were properly aware that they were in that state. - They had an unusually long outage for one particular server. - Those points combined…

I'm surprised by your risk tolerance. If I had any cloud service at this level in my stack go down for three days, I'd start shopping for an alternative. This exceeds the level of acceptability for me for even non-HA requirements. After all, if I can't trust them for this, why would I ever consider giving them my HA business? Just based on napkin math for us, this could've been a potential loss of nearly half a milli…

I think you're not exposed enough to the reality of hardware. There was no need for the host to come back online at all. I think it was a mistake of Fly.io to even attempt to do it. Just say tell the customer the host was lost and offer them a new one (with a freshly zeroed volume attached). You rent a machine, it breaks, you get a new one.

If they're sad that they lost their data, it's their fault for running on a single host with no backup. By actually performing an (apparently) difficult recovery, they reinforced their customers erroneous expectation that they are somehow responsible for the integrity of the data on any single host.

Re: Fly.io Postgres cluster down for 3 days, no word from them about it

#424

Earlier quoted context omitted.

Maybe I'll get restricted someday!

For the sake of fly.io, you should either restrict yourself and not respond or, if you can't resist, make it crystal clear, that you DO NOT represent fly.io. Your first message can and will be misunderstood and it DOES throw a poor light on fly.io. I am a paying customer of fly.io, on the Scale plan.

Please feel free to reach out directly with your concerns. I'll certainly read any email you send me.

Re: Fly.io Postgres cluster down for 3 days, no word from them about it

#425

Earlier quoted context omitted.

Not a great summary from my perspective. Here's what I got out of it: - Their free tier support depended on noticing message board activity and they didn't. - Those experiencing outages were seeing the result of deploying in a non-HA configuration. Opinions differ as to whether they were properly aware that they were in that state. - They had an unusually long outage for one particular server. - Those points combined…

Apologies for repeating myself, but: You get to a certain number of servers and the probability on any one day that some server somewhere is going to hiccup and bounce gets pretty high. That's what happened here: a single host in Sydney, one of many, had a problem. When we have an incident with a single host, we update a notification channel for people with instances on that host. They are a tiny sliver of all our us…

> over 12 hours

How much is over 12 hours? 12 hours and 10 minutes? 13 hours? 67 days?

Re: Fly.io Postgres cluster down for 3 days, no word from them about it

#426

There's a lot of bullshit in this HN thread, but here's the important takeaway: - it seems their staff were working on the issue before customers noticed it. - once paid support was emailed, it took many hours for them to respond. - it took about 20 hours for an update from them on the downed host. - they weren't updating their users that were affected about the downed host or ways to recover. - the status page was b…

(Fly.io employee here) To clarify, we communicated this incident to the personalized status page [1] of all affected customers within 30 minutes of this single host going down, and resolved the incident on the status page once it was resolved ~47h later. Here's the timeline (UTC): - 2023-07-17 16:19 - host goes down - 2023-07-17 16:49 - issue posted to personalized status page - 2023-07-19 15:00 - host is fixed - 202…

Dude. I don't sit at home refreshing status pages. Send me an e-mail.

That's how other [useful] providers notify their customers that one of their hosts went down unexpectedly. Linode will send me 6 emails when they need to reboot something. Even Oracle sends me notices about network blips. I believe I've gotten one from AWS, but I also know sometimes their gear gets stuck in a bad state and I didn't get a notification, which was super annoying because it took forever to figure out it was AWS's faulty state.

Re: Fly.io Postgres cluster down for 3 days, no word from them about it

#427
After a while the direct link to Google Search's cache will no longer work but it appears the the original link is now accessible: https://community.fly.io/t/service-interruption-cant-destroy...

Anyway, here's an archive link for future visitors: https://archive.is/7lSJA

Re: Fly.io Postgres cluster down for 3 days, no word from them about it

#428
post #400

Earlier quoted context omitted.

I see. Have you considered eliminating this configuration from your offering? It sounds like the terminology could confuse people, and it may be the case that they're assuming that a host isn't really what it is (a single host). This kind of thing is difficult for those seeking to build managed services, because I think people expect you to provide offerings that can't harm them when the cause is related to the servi…

Plenty of people would rather take downtime than pay for redundancy, for example for a test database. AWS RDS lets you spin up a RDS instance that costs 3x less and regularly has downtime (the 'single-az' one), quite similar to this. Anyone who's used servers before knows "A single instance" is the same as "sometimes you might have downtime". Computers aren't magic, everyone from heroku (you must have multiple dynos…

Single-AZ i not single-host though, and while a single AZ can go down for major events, it doesn't break because a single piece of hardware failed.

Re: Fly.io Postgres cluster down for 3 days, no word from them about it

#429

There's a lot of bullshit in this HN thread, but here's the important takeaway: - it seems their staff were working on the issue before customers noticed it. - once paid support was emailed, it took many hours for them to respond. - it took about 20 hours for an update from them on the downed host. - they weren't updating their users that were affected about the downed host or ways to recover. - the status page was b…

Haha, imagine what the AWS status page would look like if they had to update their global status page anytime a single host would go down in any region.

Fly.io messed up, they didn't want to be a Heroku clone, but their marketing and their polished user experience design made it seem like they would be one anyway.

And as a reward now they have to deal with bottom of the barrel Heroku users that manage to do major damage to their brand whenever a single host goes down. Who would have predicted that corporate risk?

Re: Fly.io Postgres cluster down for 3 days, no word from them about it

#430
post #428

Earlier quoted context omitted.

Plenty of people would rather take downtime than pay for redundancy, for example for a test database. AWS RDS lets you spin up a RDS instance that costs 3x less and regularly has downtime (the 'single-az' one), quite similar to this. Anyone who's used servers before knows "A single instance" is the same as "sometimes you might have downtime". Computers aren't magic, everyone from heroku (you must have multiple dynos…

Single-AZ i not single-host though, and while a single AZ can go down for major events, it doesn't break because a single piece of hardware failed.

Sure, but isn't this more about risk tolerance at this point and how much your customers care about? Where the responsibility should be on customer's end. Running on EBS/RDS doesn't guarantee you won't lose data. If you care about it, you enable backups and test recovery.

Just because some customers are less fault tolerant than others, doesn't mean we shouldn't offer those options where people don't have the same requirements or are willing to work around it.

Post reply on HN