Live data from Hacker News

Fly.io outage – resolved

status.flyio.net

121–130 of 287 posts

Re: Fly.io outage – resolved

#121

Earlier quoted context omitted.

The tech is impressive and the pricing is attractive which is why we use them. I just wish there was less black magic.

I don’t always agree with @tptacek on social/political issues, and I don’t always agree with @xe on the direction of Nix, but these are legends on the technical side of things. And they’re trying to build an equitable relationship between the user of cloud services and the provider, not fund a private space program. If I were in the market for cloud services I’d highly prize a long-term relationship on mutual benefit…

FWIW Xe was let go from Fly earlier this year during a round of layoffs.

Re: Fly.io outage – resolved

#122
post #92

fly.io publishes their post-mortems here: https://fly.io/infra-log/ The last post-mortem they wrote is very interesting and full of details. Basically back in 2016 the heart or keystone component of fly.io production infrastructure was called consul, which is a highly secure TLS server that tracks shared state and it requires that both the server certificate and the client certificate be authenticated. Since it was c…

On that Consul outage, Fly Infra concludes, "The moral of the story is, no more half-measures."

On their careers page [1], the Fly team goes, "We're not big believers in tech debt."

As an outsider, reads like a cacophony of contradictions?

[1] https://fly.io/docs/hiring/working/#we-re-ruthless-about-doi...

Re: Fly.io outage – resolved

#123

Earlier quoted context omitted.

I don’t always agree with @tptacek on social/political issues, and I don’t always agree with @xe on the direction of Nix, but these are legends on the technical side of things. And they’re trying to build an equitable relationship between the user of cloud services and the provider, not fund a private space program. If I were in the market for cloud services I’d highly prize a long-term relationship on mutual benefit…

FWIW Xe was let go from Fly earlier this year during a round of layoffs.

Unfortunate. Xe rocks.

Re: Fly.io outage – resolved

#124
post #110

Kinda funny that they've named their global state store "Corrosion"... not really a word I'd associate with stability and persistence.

I stored important data in mnesia, so who would I be to talk. :p

amnesia means forget, so mnesia means remember, I would guess?

Re: Fly.io outage – resolved

#125

Color me not surprised. My few interactions with people there just gave off the impression of them being in a bit over their heads. I don't know how well that translated to their actual ops, but it's difficult to not connect the two when they continue to have major outage after major outage for a product that 'should' be their customer's bedrock upon which they build everything else.

[deleted]

Re: Fly.io outage – resolved

#126

Earlier quoted context omitted.

The tech is impressive and the pricing is attractive which is why we use them. I just wish there was less black magic.

I don’t always agree with @tptacek on social/political issues, and I don’t always agree with @xe on the direction of Nix, but these are legends on the technical side of things. And they’re trying to build an equitable relationship between the user of cloud services and the provider, not fund a private space program. If I were in the market for cloud services I’d highly prize a long-term relationship on mutual benefit…

Xe here. As a sibling comment said, I didn't survive layoffs. If you're looking for someone like me, I'm on the market!

Re: Fly.io outage – resolved

#127
post #58

Earlier quoted context omitted.

It's still 99.99+% SLA? Would you really pay 100% more for <0.01% more uptime?

I think what a lot of people fail to understand is that there are certain categories of apps that simply “can never go down” Examples include basically any PaaS, IaaS, or any company that provides a mission-critical service to another company (B2B SaaS). If you run a basic B2C CRUD app, maybe it’s not a big deal if you service goes down for 5 minutes. Unfortunately there are quite a few categories of companies where…

My experience with very large scale B2B SaaS and PaaS has been that customers like to get money, if allowed by contract, by complaining about outages, but that overall, B2B SaaS is actually very forgiving.

Most B2B SaaS solutions have very long sales cycles and a high total cost to implement, so there is a lot of inertia to switching that “a few annoying hours of downtime a year” isn’t going to cover. Also, the metric that will drive churn isn’t actually zero downtime, it’s “nearest competitor’s downtime,” which is usually a very different number.

Re: Fly.io outage – resolved

#128
post #106
post #91

Recurring pattern I notice is outages tend to occur the week of major holidays in US. - MS 365/Teams/Exchange had a blip in the morning - Fly.io with complete outage - then a handful of sites and services impacted due to those outages Usually advocate against “change freezes” but I think a change freeze around major holidays makes sense. Give all teams a recharge/pause/whatever. Don’t put too much pressure on the B-s…

Then you just get devs rushing out changes before the freeze…

and stampeding changes in after the thaw, also leading to downtime. so it depends on the org, but doing a freeze is still reasonable policy. Downtime on December 15th is less expensive than on black Friday or cyber Monday for most retailers, so it's just a business decision at that point.

Re: Fly.io outage – resolved

#129
post #91

Recurring pattern I notice is outages tend to occur the week of major holidays in US. - MS 365/Teams/Exchange had a blip in the morning - Fly.io with complete outage - then a handful of sites and services impacted due to those outages Usually advocate against “change freezes” but I think a change freeze around major holidays makes sense. Give all teams a recharge/pause/whatever. Don’t put too much pressure on the B-s…

What do "Freezes" mean? Like, do you stop renewing your certificates? Do you stop taking in security updates for your software? Sure maybe "unnecessary" changes, but the line gets very gray very fast.

It's not very grey, prod becomes as if you told everyone but your ops team to go home and then sent your ops team on a cruise with pagers. If it's not important enough to merit interrupting their vacation you don't do it.

Re: Fly.io outage – resolved

#130
post #126

Earlier quoted context omitted.

I don’t always agree with @tptacek on social/political issues, and I don’t always agree with @xe on the direction of Nix, but these are legends on the technical side of things. And they’re trying to build an equitable relationship between the user of cloud services and the provider, not fund a private space program. If I were in the market for cloud services I’d highly prize a long-term relationship on mutual benefit…

Xe here. As a sibling comment said, I didn't survive layoffs. If you're looking for someone like me, I'm on the market!

Hiring people is above my pay grade, but I can vouch to my lords and masters and anyone else who cares what I think that a legend is up for grabs.

b7r6@b7r6.net

Post reply on HN