Live data from Hacker News

Reliability: It’s not great

community.fly.io

291–300 of 476 posts

Re: Reliability: It’s not great

#291

Almost half of the issues are caused by their use of HashiCorp products. As someone that has started tons of Consul clusters, analyzed tons of Terraform states, developed providers and wrote a HCL parser, I must say this: HashiCorp built a brand of consistent design & docs, security, strict configuration, distributed-algos-made-approachable... but at its core, it's a very fragile ecosystem. The only benefit of HashiC…

We are asking to HashiCorp products to do things they were not designed to do, in configurations that they don't expect to be deployed in. Take a step back, and the idea of a single global namespace bound up with Raft consistency for a fleet deployed in dozens of regions, providing near-real-time state propagation, is just not at all reasonable. Our state propagation needs are much closer to those of a routing protoc…

Are there any plans to make Corrosion open source? Or are you able to talk at all about the technologies/patterns used to create it? I feel like service discovery is still ripe for disruption

Re: Reliability: It’s not great

#292

Earlier quoted context omitted.

We are asking to HashiCorp products to do things they were not designed to do, in configurations that they don't expect to be deployed in. Take a step back, and the idea of a single global namespace bound up with Raft consistency for a fleet deployed in dozens of regions, providing near-real-time state propagation, is just not at all reasonable. Our state propagation needs are much closer to those of a routing protoc…

Are there any plans to make Corrosion open source? Or are you able to talk at all about the technologies/patterns used to create it? I feel like service discovery is still ripe for disruption

Yeah, we'll for sure talk about it more some other time. Mostly today we want to talk about how we were sucking ass at customer comms.

Re: Reliability: It’s not great

#293
post #159

Fundamentally I think some of the problems come down to the difference between what Fly set out to build and what the market currently want. Fly (to my understanding) at its core is about edge compute. That is where they started and what the team are most excited about developing. It's a brilliant idea, they have the skills and expertise. They are going to be successful at it. However, at the same time the market is…

This is, indeed, the exciting part. As Heroku fans, we never really felt like it needed a replacement. And if it did, it seemed like Render was the natural Heroku v.next. One thing we've noticed, though, is that people do actually want Heroku but close to users. It's not exactly edge compute. In some cases, it's "Heroku in Tokyo". In others it's "Heroku, but running in all the english speaking regions". I think the t…

It sounds like you built the right product on the wrong technology stack, at the lower levels. For example, I have never heard a Nomad success story, but this might be colored by interviewing engineers desperate to escape it.

Something like linkerd on Kubernetes would be stronger, I suspect. But I don't know the exact nature of your problems.

Re: Reliability: It’s not great

#294
post #118
post #95

Earlier quoted context omitted.

I'm not sure why the cynicism around their candor. Do you think it's not genuine just because it was posted by a company employee? Your post implies corporate messaging is bad. And anything posted by a company—or at least I don't know where you draw the line—can be considered corporate messaging. Am I just reading too much into your phrasing?

It's strategic messaging. It can't be genuine, because of what it is. The benefit they get is publicity and damage control, and as you can tell by the many responses here, it buys them time because many developers are willing to give them the benefit of the doubt. Companies that engage in this kind of candor are careful not to disclose those things that would really hurt their business. Those things are still kept se…

Yeah I agree with you. Corporate messaging is still corporate messaging. If you go the candor route, it's a tactic of many you can employ.

Re: Reliability: It’s not great

#295
post #5

I'm not a user of Fly.io. I can't help but notice how remarkable the effect of open communication on potential end users like me. I remember reading about their reliability problems on HN some time ago. That biased my view of the company. After reading this, the open communication and transparency restored my trust in them, and would make them again a potential candidate for future projects. Because now I know that t…

I’m not [directly, at present] in the market for their product but the very frank and real introduction basically moved Fly from my “I think I’ve heard of that but I’ll need my memory jogged for any further recognition” category to my “potentially good stuff to reference when relevant” category.

This kind of frank and human communication is vulnerable, but it’s good for establishing credibility… with me at least!

Re: Reliability: It’s not great

#296
post #210

Earlier quoted context omitted.

Yeah, distributed systems at the global scale are very very difficult - at least with the Heroku style problem, you'd be looking at scaling in a single datacenter I think - deployments to multiple datacenters wouldn't share dependencies. I do wonder however if they'd be better off using less l33t tech - do almost everything on Postgres vs consul and vault, etc. Scaling, failover, consistency, etc is a more well-known…

> I do wonder however if they'd be better off using less l33t tech - do almost everything on Postgres vs consul and vault, etc. Scaling, failover, consistency, etc is a more well-known problem and there are a lot of people who've ran other DBs at tremendous scale than the alternatives. In my experience people who ran Postgres distributed across a WAN tended to use obscure third-party plugins at best, more often a pil…

Yeah, point taken, I wasn't thinking to cluster across the WAN - more like an api wrapping postgres in a single DC. But you pay the price of read latency I guess... it's a hard problem no doubt.

Re: Reliability: It’s not great

#297

Earlier quoted context omitted.

I am willing to pay a little extra for a nice dev/ops experience and simple/easy solutions that doesn't require spending days reading docs and diving into dashboards with thousands of options. Usually this results in me jumping on new platforms and then abandoning them once they add too much complexity.

I suspect, in general, acceptability (or desire) for complexity in the cloud solution, and budget are positively correlated in customers.

The ridiculously overwhelming complexity is stickiness.

Think it’s bad to potentially technically move your solution from $CLOUD vendor? Wait until you turn around and realize you have at least one full time hire who’s entire role is “$BIGCLOUD Certified Architect” (or whatever) and your entire dev staff was also at least partially selected for experience with the preferred cloud vendor. At any kind of scale you have massive amounts of tooling, technical debt, and institutional knowledge built around the cloud provider of choice.

Then there’s all of the legal, actually understanding billing (pretty much impossible but you’re probably close by now), etc elsewhere in the org. At this point you’ve probably utilized an outside service/consultant or two from the entire cottage industry that has sprung up to plug holes in/augment your cloud provider of choice.

After realizing their cloud spend has ballooned well beyond what they ever anticipated plenty of orgs get far enough to investigate leaving before they realize all of this. Most decide to suck it up and keep paying, or try to somehow negotiate or optimize spend down further.

Cloud platforms are a true masterclass in customer stickiness and retention - to the Oracle and Microsoft level (who also operate clouds).

It’s interesting here on HN because while MS and Oracle are bashed for these practices AWS and GCP (for the most part) are pretty beloved for what are really the same practices.

Re: Reliability: It’s not great

#298
post #62

Earlier quoted context omitted.

The CloudFlare folks wrote a good blog post on how they are seeing their customers use Edge compute — latency is far down on the list: https://blog.cloudflare.com/cloudflare-workers-serverless-we...

The US CLOUD Act means a EU customer cannot use a US cloud provider to host PII, even if the server itself is physically in the EU, because US law will still compel the provider to yield the data to US authorities. The European Commission is trying to paper over the cracks with a fig leaf of judicial review, but it's only a matter of time until a Schrems III decision from the CJEU invalidates that polite fiction.

Not exactly related to the OP, but: I think I speak for a large number of folks when I say that we don't care. The EU keeps passing all sorts of absurd laws that require dedicated auditors to comply with. It's just not going to happen. If they decide to actively enforce these things, they'll just isolate themselves from the rest of the world.

Re: Reliability: It’s not great

#299

Earlier quoted context omitted.

I have never seen a company without Google Search, Google Chrome, AWS, Microsoft 360 and the lot. Which alternatives are they based on?

Those would not contain PII from your users though, unless you have terrible policies about copying personal information in random Google Docs.

Companies have to guess what is PII and what is not, the EU have no idea (other than they know which companies they want to punish)

Re: Reliability: It’s not great

#300

Well, I feel for them. Scaling up is a bitch. I've been lucky, in the past, but a lot of that, is because I have "overengineered," and the tools/frameworks have advanced to meet the new demand. I am in the middle of a complete, bottom-to-top rewrite of the app we've been developing for the last couple of years. It's going great, but making this leap was a fraught decision. It's mainly, so I wouldn't have to write a p…

Don’t know about expected usage but 10K rows seems like low number to me. If you mean 10K req/s then perf starts to matter but usually it’s SQL queries that fail first if you have 100K+ rows. In general good caching solves most of stuff + read/write separation. That’s all from my poor experience. What I mean is that these scaling problems don’t have much to do with app logic but having nice core is good so :thumbsup:

10K is a rounding error, to most folks, around here.

The app (and the demographic that it Serves) are very security and privacy-conscious, so security and privacy are the main coefficients. Most folks would be disappointed in how few features the app boasts, as each feature is a potential compromise. I'm glad to have low usage, as a result.

I wrote the backends, as well as the frontend, and have avoided third-party dependencies, all around.

It seems one of the first casualties of fast scaling is security.

I really didn't want to take any chances. An overly-complex architecture is begging for security compromises.

Post reply on HN