Nice to see a provider being transparent about KV cache quantisation. I've been suspecting that some providers do this silently whilst heavily promoting their unquantised weights, even though KV quantisation can degrade quality more than weight quantisation. However, I wish their testing were more detailed. Firstly, some model families are more sensitive to KV quantisation than others (only Kimi K2.6 was tested). Sec…
vLLM's study also concluded that "FP8 can deliver meaningful latency and capacity gains with small or negligible accuracy loss". Their benchmarks include LiveCodeBench 6. https://vllm-project.github.io/2026/04/22/fp8-kvcache.html
Smaller, faster, safer: running Kimi and GLM at scale
31–40 of 74 posts
Re: Smaller, faster, safer: running Kimi and GLM at scale
#32> View pricing in the Cloudflare dashboard ↗ Why… I wanted to see if it’s worth it to use cloudflare’s endpoint but I can’t even see the pricing
I'm not sure how accurate this is, but there is pricing here: https://openrouter.ai/provider/cloudflare
Re: Smaller, faster, safer: running Kimi and GLM at scale
#33Earlier quoted context omitted.
Drives me nuts that comments are held to a higher standard than submissions. HN is for conversation between humans[1] (about AI generated blogspam, apparently) 1. https://news.ycombinator.com/newsguidelines.html
Meta recently added a filter as a requirement for posts on Facebook. if it was ai generated, you are required to check off a box for that on your posts. I've been asking for that for some time.
Re: Smaller, faster, safer: running Kimi and GLM at scale
#34I was interested in reading this until my slop detector went off at the paragraph starting with “It's worth being precise about where the benefit comes from, because it isn't raw speed.” I love AI, but I really hate reading it.
Re: Smaller, faster, safer: running Kimi and GLM at scale
#35I was interested in reading this until my slop detector went off at the paragraph starting with “It's worth being precise about where the benefit comes from, because it isn't raw speed.” I love AI, but I really hate reading it.
I have had to stop commenting this because it would end up on 50% of the posts here. I really wish we could flag prose as ai-generated on here and just filter it out.
Re: Smaller, faster, safer: running Kimi and GLM at scale
#36> If squeezing the best open models onto GPUs and serving them to millions of developers sounds like your kind of problem, come work with us. What is the typical job title and/or skillset for this?
It can be any one of many jobs depending on how high close to the metal one's focus is, but the highest headcount role is usually SRE/Infra/Ops with GPU knowledge sprinkled on top. That is to say Linux sysadmin, networking, fleet management, scaling, incident troubleshooting, etc.
Re: Smaller, faster, safer: running Kimi and GLM at scale
#37Re: Smaller, faster, safer: running Kimi and GLM at scale
#38I was interested in reading this until my slop detector went off at the paragraph starting with “It's worth being precise about where the benefit comes from, because it isn't raw speed.” I love AI, but I really hate reading it.
Re: Smaller, faster, safer: running Kimi and GLM at scale
#39Earlier quoted context omitted.
I'm not sure how accurate this is, but there is pricing here: https://openrouter.ai/provider/cloudflare
So they don't even support K3? What's the point. K2.7 Code is practically free already
Input tokens (per 1M)$3.00 Cached input tokens (per 1M)$0.30 Output tokens (per 1M)$15.00
Re: Smaller, faster, safer: running Kimi and GLM at scale
#40> If squeezing the best open models onto GPUs and serving them to millions of developers sounds like your kind of problem, come work with us. What is the typical job title and/or skillset for this?
I've seen this called MLOps.
Applied ML Research Engineer or something maybe, not that I have ever seen that title. Maybe just catchall ML Engineer...