Live data from Hacker News

Skipping 90% of KV dequant work speeds up LLM decode by 22%

github.com

1–2 of 2 posts