Live data from Hacker News

Show HN: Proxima serves 4x more requests with no hardware change on vLLM

github.com

1–2 of 2 posts

Show HN: Proxima serves 4x more requests with no hardware change on vLLM

#1
hey everyone, i decided to make a vLLM plugin that implements the Star-KV paper. the results are quiet promising with a decode kernel thats faster than FA2 in higher batch sizes. would love any thoughts and recommendations

Show HN: Proxima serves 4x more requests with no hardware change on vLLM
github.com