Why this is not a PR for vLLM ?
Re: KVarN: Native vLLM backend for KV-cache quantization by Huawei
#11Last I heard, vLLM was backed by a company that has raised $150m in seed funding. I'm sure they've got the resources to port it.
11–20 of 21 posts
Why this is not a PR for vLLM ?
Better performance than TQ and better quality than FP16? Am I reading this right??
Why this is not a PR for vLLM ?
... and it's on llama.cpp that to this guy! https://www.reddit.com/r/LocalLLaMA/comments/1txlhxu/i_imple...