Live data from Hacker News

Llama.cpp speculative sampling: 2x faster inference for large models

github.com

1–2 of 2 posts