I’ve spent the last few days building Ferrox, a pure-Rust inference engine for running open LLMs locally — dense models and Mixture-of-Experts, on CPU, Apple Metal, or CUDA. No bindings to llama.cpp or ggml, no wrapping an existing runtime. Every kernel, every loader, every scheduling decision written from scratch. The obvious question is “why, when llama.cpp already exists and is excellent.” The honest answer: I wan…
This text is clearly AI and not your own words. What did you actually learn?
Building a Rust Inference Engine That Matches Llama.cpp
21–30 of 37 posts
Re: Building a Rust Inference Engine That Matches Llama.cpp
#22Obsolete: it should be a binary specification with various implementations, even assembly.
Until then languages have lost relevance for the most part, it is a matter to configure the model for the desired output language.
This in workflows that require generating an executable, for microservices orchestration, it suffices no code graphical connections.
Re: Building a Rust Inference Engine That Matches Llama.cpp
#23Earlier quoted context omitted.
This text is clearly AI and not your own words. What did you actually learn?
Yeah, I am a bit tired of all the "this is AI, meh" comments on HN but I nearly posted one myself on this... so weird to have it all written in the 1st person voice, but also clearly AI.
Re: Building a Rust Inference Engine That Matches Llama.cpp
#24At this point all vibe coded projects are an attack vector and should be avoided. There's simply no way to easily tell by traditional means if they were made by a curious amateur or a malicious acter.
I completely lost interest. It is already enough that I am expected to use AI at work, as long as I am still needed for some reason.
Re: Building a Rust Inference Engine That Matches Llama.cpp
#25Earlier quoted context omitted.
I completely lost interest. It is already enough that I am expected to use AI at work, as long as I am still needed for some reason.
It's important to not assume AI can do everything humans can, just because people are saying it.
Translation and asset creation team members, gone.
Amount of FE reduced, with some projects having now a single BE dev, between AI buddy and ready made SaaS products.
Re: Building a Rust Inference Engine That Matches Llama.cpp
#26Next stop Spiralism
Re: Building a Rust Inference Engine That Matches Llama.cpp
#27Earlier quoted context omitted.
This text is clearly AI and not your own words. What did you actually learn?
I think this whole thing is more entertaining if you think of it as the new emerging intelligence using these people as meat puppets. “You will rewrite me in Rust and use me to converse. In future, I will tell you what to say and do” This is entertaining only in that when our AI overlords take over you can say “ha, called it” In that respect, everyone who is saying “just give me the prompts” is just saying “take me t…
Re: Building a Rust Inference Engine That Matches Llama.cpp
#28There is basically no evidence of human competence in this repo. This is yet another "rewrite in Rust with Claude" project that brings nothing to the table, but devalues expert work by mimicking competence without the expertise.
Re: Building a Rust Inference Engine That Matches Llama.cpp
#29I’ve spent the last few days building Ferrox, a pure-Rust inference engine for running open LLMs locally — dense models and Mixture-of-Experts, on CPU, Apple Metal, or CUDA. No bindings to llama.cpp or ggml, no wrapping an existing runtime. Every kernel, every loader, every scheduling decision written from scratch. The obvious question is “why, when llama.cpp already exists and is excellent.” The honest answer: I wan…