> We are working to build a new Consul cluster with 10x the RAM. Oh boy. I wouldn’t wanna be the people doing this. Working with infrastructure is hard . Doing it under tight SLAs? Ugh. I really hope the people working on this are being well supported.
You wouldn't necessarily know this from the outside, but we have _exceptional_ internal support when things go sideways. This is relatively new, up until about two months ago most incidents were run by 1.5 people. We had 7 people working this one today.
Fly.io Status – Consul cluster outage
51–60 of 123 posts
Re: Fly.io Status – Consul cluster outage
#52This has been a rough week, and I'm sorry we broke peoples' apps. We had a big Nomad outage on Monday, and then a suspiciously similar Consul outage today. Both tipped over faster than we could detect and mitigate, and we ended up having to do serious surgery to build entirely new Consul/Nomad clusters. There's nothing to brag about here, I just wanted to let y'all know we're listening (even when things aren't on the…
We had mysterious consul outages (and related nomad outages) causing us to never deploy our new hashicorp stack to production. Shame cuz we were excited about our nomad+consul+vault setup and invested a lot of money into building it. But just didn’t have the time or enough depth of expertise to babysit it.
Still love using Fly, please add static assets hosting/CDN.
Re: Fly.io Status – Consul cluster outage
#53They seem to have a lot of issues with Consul, is it the design of Consul or the way they use it that is the problem?
The latter (they openly admit as such)
I don't have a relationship with Hashicorp, and have tried using Consul. Everything about it is amazing in theory, but you might need a few years of experience with kube, consul, go, and maybe even the hashicorp stack to even begin debugging when things don't work as advertised.
I still think my company is going to take another stab at consul in the future, because we do need service discovery. But they're advertising a solution to an incredibly hard problem with a shit ton of variations in network topology and infra that it should (theoretically) work on. I imagine if you stay on the happy path everything works out just fine with Consul (even then, maybe only most of the time). The problem is that they don't spell out what the happy path is, and that all the other knobs they expose off to the side are actually down paths beleagured by dragons.
Re: Fly.io Status – Consul cluster outage
#54They seem to have a lot of issues with Consul, is it the design of Consul or the way they use it that is the problem?
Re: Fly.io Status – Consul cluster outage
#55Earlier quoted context omitted.
We had mysterious consul outages (and related nomad outages) causing us to never deploy our new hashicorp stack to production. Shame cuz we were excited about our nomad+consul+vault setup and invested a lot of money into building it. But just didn’t have the time or enough depth of expertise to babysit it.
From my experience with the Hashi stack, I don't think it's a coincidence that Fly has a lot of downtime and are a major Hashi user. Terraform makes excellent bait though. Still love using Fly, please add static assets hosting/CDN.
Re: Fly.io Status – Consul cluster outage
#56Earlier quoted context omitted.
Are they really building everything in hard mode or do they just have a bad architecture?
Yes. Both. This outage was caused by a bad architectural decision. We had an incident a few weeks ago caused by "hard mode".
Re: Fly.io Status – Consul cluster outage
#57Earlier quoted context omitted.
Same. I’m so disappointed because I’ve been rooting for them. We were close to a major deployment/migration (well, major as is mid four figures per month, not major like Google) but they were removed from the decision set. It would not have been responsible to bet on them at this time. I hope they get this sorted - they’re really good folks!
Thank you! I'm both sorry it didn't work out (because $$$$) and also glad we didn't create any agony for you. Someday, we hope to create mild irritation for you, though, if we can.
Re: Fly.io Status – Consul cluster outage
#58At this point I'm not sure why one wouldn't use something like Hetzner and slap Coolify or Dokku or something else on it.
[1] We're using so little infra at present that we're within their free usage tier. However, I want to clarify that this isn't because we aren't willing to pay, we specifically want to pay for reliable managed offerings. That's actually the entire point! If Fly.io can deliver on their vision, we'd gladly be billed at 100x the current usage rates.
Re: Fly.io Status – Consul cluster outage
#59Earlier quoted context omitted.
Are they really building everything in hard mode or do they just have a bad architecture?
Yes. Both. This outage was caused by a bad architectural decision. We had an incident a few weeks ago caused by "hard mode".
Re: Fly.io Status – Consul cluster outage
#60Earlier quoted context omitted.
From my experience with the Hashi stack, I don't think it's a coincidence that Fly has a lot of downtime and are a major Hashi user. Terraform makes excellent bait though. Still love using Fly, please add static assets hosting/CDN.
How is this possible? How is consul not self—healing? It just seems so brittle in a way even database clusters aren’t.