KVarN: Native vLLM backend for KV-cache quantization by Huawei
1–10 of 21 posts
Re: KVarN: Native vLLM backend for KV-cache quantization by Huawei
#2Re: KVarN: Native vLLM backend for KV-cache quantization by Huawei
#3Am I reading this right??
Re: KVarN: Native vLLM backend for KV-cache quantization by Huawei
#4Why this is not a PR for vLLM ?
edit: It might not be clear that it is based on vLLM 0.22, which is the current version: https://github.com/huawei-csl/KVarN/commit/d6290e99098d7426d.... All you have to do is create a diff off it; it's fairly straightforward.
Re: KVarN: Native vLLM backend for KV-cache quantization by Huawei
#5Why this is not a PR for vLLM ?
It's the output of a research paper; the authors are not trying to build up vLLM, and they probably have no incentive to do so. You can submit a PR, though! It's easier now while the divergence is low, so don't wait. Since there are six authors, I bet you could get help with the inevitable review chores if you just take the step of creating the PR. edit: It might not be clear that it is based on vLLM 0.22, which is t…
Re: KVarN: Native vLLM backend for KV-cache quantization by Huawei
#6Better performance than TQ and better quality than FP16? Am I reading this right??
Re: KVarN: Native vLLM backend for KV-cache quantization by Huawei
#7Better performance than TQ and better quality than FP16? Am I reading this right??
Re: KVarN: Native vLLM backend for KV-cache quantization by Huawei
#8Better performance than TQ and better quality than FP16? Am I reading this right??
Re: KVarN: Native vLLM backend for KV-cache quantization by Huawei
#9Re: KVarN: Native vLLM backend for KV-cache quantization by Huawei
#10Why this is not a PR for vLLM ?