Four day old repo with 92K lines added per day, okay... There is basically no evidence of human competence in this repo. This is yet another "rewrite in Rust with Claude" project that brings nothing to the table, but devalues expert work by mimicking competence without the expertise.
Building a Rust Inference Engine That Matches Llama.cpp
31–36 of 36 posts
Re: Building a Rust Inference Engine That Matches Llama.cpp
#32Re: Building a Rust Inference Engine That Matches Llama.cpp
#33Earlier quoted context omitted.
It's important to not assume AI can do everything humans can, just because people are saying it.
I have already seen the real impact on delivery team sizes, no need to hear to what people are saying about what AI can do or not. Translation and asset creation team members, gone. Amount of FE reduced, with some projects having now a single BE dev, between AI buddy and ready made SaaS products.
Re: Building a Rust Inference Engine That Matches Llama.cpp
#34Earlier quoted context omitted.
I have already seen the real impact on delivery team sizes, no need to hear to what people are saying about what AI can do or not. Translation and asset creation team members, gone. Amount of FE reduced, with some projects having now a single BE dev, between AI buddy and ready made SaaS products.
well just a few years ago we were all fullstack anyway. Back to it.
Re: Building a Rust Inference Engine That Matches Llama.cpp
#35Wonderful... I'm so happy to see a Rust version of llama.cpp. The true value of this will be proven over time with wide use and as PRs are merged. Do you have a feel if you'll try to drive this to stay feature parity with llama.cpp, or are you willing to diverge with new features like NVME/SSD MoE weight streaming etc.
What I actually want is MoE on machines that can't fit the model in VRAM, and specifically expert-level residency instead of layer offload: track which experts get hit during decode, keep those resident, evict the rest. Doing that well needs the router, the KV cache and the memory manager to be designed together, which is about the only good reason to write a runtime from scratch.
Yes, let’s see! You are welcome to contribute if you like!
Re: Building a Rust Inference Engine That Matches Llama.cpp
#36I’ve spent the last few days building Ferrox, a pure-Rust inference engine for running open LLMs locally — dense models and Mixture-of-Experts, on CPU, Apple Metal, or CUDA. No bindings to llama.cpp or ggml, no wrapping an existing runtime. Every kernel, every loader, every scheduling decision written from scratch. The obvious question is “why, when llama.cpp already exists and is excellent.” The honest answer: I wan…
I have the skills, qualifications and experience to do this; ( see https://www.credly.com/users/antonello-fratepietro ) by using AI (I monitor what it does), I’m able to carry out a project like this! I’m a coder, not a blogger; it’s only natural that I use AI, just like everyone else, to write technical texts. I have over 20 years’ experience and I’m not ashamed to admit that I use Claude, Cursor and so on. The prob…