"Small models" will always outperform as they are deterministic (or closer to it). This was realized in 2023 already: https://newsletter.semianalysis.com/p/google-we-have-no-moat... "Less is best" is not a new realization. The concept exists across contexts. Music described as "overplayed". Prose described as verbose. We just went through an era of compute that chanted "break down your monoliths". NPM ecosystem being…
The "no moat" memo you linked was about open source catching up to closed models through fine-tuning, not about small models outperforming large ones.
I'm also not sure what "skeletonized down to opcodes" or "geometry for text as bytecode patterns" means in the context of neural networks. Model compression is a real field (quantization, distillation, pruning) but none of it works the way you're describing here.