Live data from Hacker News

Re-quantizing a local LLM 14x faster by skipping the tensors that didn't change

andreaborio.substack.com

1–2 of 2 posts