https://arxiv.org/abs/2607.14431
Making models smarter and cheaper at the same time
1–2 of 2 posts
Re: Making models smarter and cheaper at the same time
#2Byte-exact KV grafting stores verified reasoning on disk. Replaying it lets a frozen 12B LLM beat 31B models at 8,700x less energy.