Llama.cpp speculative sampling: 2x faster inference for large models
1–2 of 2 posts
Re: Llama.cpp speculative sampling: 2x faster inference for large models
#2Small caveat: this is only true for generating text with simple grammar, like code. For human text this doesn't work so well.