Simple, zero overhead way to compress model, KV cache via Low-Rank Decomposition
jeffreywong20.github.io