Live data from Hacker News

Smaller, faster, safer: running Kimi and GLM at scale

blog.cloudflare.com

31–40 of 74 posts

Re: Smaller, faster, safer: running Kimi and GLM at scale

#31
post #15
post #4

Nice to see a provider being transparent about KV cache quantisation. I've been suspecting that some providers do this silently whilst heavily promoting their unquantised weights, even though KV quantisation can degrade quality more than weight quantisation. However, I wish their testing were more detailed. Firstly, some model families are more sensitive to KV quantisation than others (only Kimi K2.6 was tested). Sec…

vLLM's study also concluded that "FP8 can deliver meaningful latency and capacity gains with small or negligible accuracy loss". Their benchmarks include LiveCodeBench 6. https://vllm-project.github.io/2026/04/22/fp8-kvcache.html

[deleted]

Re: Smaller, faster, safer: running Kimi and GLM at scale

#32
post #27

> View pricing in the Cloudflare dashboard ↗ Why… I wanted to see if it’s worth it to use cloudflare’s endpoint but I can’t even see the pricing

I'm not sure how accurate this is, but there is pricing here: https://openrouter.ai/provider/cloudflare

So they don't even support K3? What's the point. K2.7 Code is practically free already

Re: Smaller, faster, safer: running Kimi and GLM at scale

#33
post #24

Earlier quoted context omitted.

Drives me nuts that comments are held to a higher standard than submissions. HN is for conversation between humans[1] (about AI generated blogspam, apparently) 1. https://news.ycombinator.com/newsguidelines.html

Meta recently added a filter as a requirement for posts on Facebook. if it was ai generated, you are required to check off a box for that on your posts. I've been asking for that for some time.

This is only to help them filter out AI generations for their own training data. There is no way you can "block" all "ai generated content" from your view.

Re: Smaller, faster, safer: running Kimi and GLM at scale

#35
post #3

I was interested in reading this until my slop detector went off at the paragraph starting with “It's worth being precise about where the benefit comes from, because it isn't raw speed.” I love AI, but I really hate reading it.

I have had to stop commenting this because it would end up on 50% of the posts here. I really wish we could flag prose as ai-generated on here and just filter it out.

Recently, I wrote a small script for myself that scrapes the comments to see if anyone already said it's AI, and if so, grays out the article...

Re: Smaller, faster, safer: running Kimi and GLM at scale

#36

> If squeezing the best open models onto GPUs and serving them to millions of developers sounds like your kind of problem, come work with us. What is the typical job title and/or skillset for this?

What is the typical job title and/or skillset for this?

It can be any one of many jobs depending on how high close to the metal one's focus is, but the highest headcount role is usually SRE/Infra/Ops with GPU knowledge sprinkled on top. That is to say Linux sysadmin, networking, fleet management, scaling, incident troubleshooting, etc.

Re: Smaller, faster, safer: running Kimi and GLM at scale

#38

I was interested in reading this until my slop detector went off at the paragraph starting with “It's worth being precise about where the benefit comes from, because it isn't raw speed.” I love AI, but I really hate reading it.

I hope you can get over it, because before long literally everything will be written by AI. Millions of people write with this every day.

Re: Smaller, faster, safer: running Kimi and GLM at scale

#39
post #32
post #27

Earlier quoted context omitted.

I'm not sure how accurate this is, but there is pricing here: https://openrouter.ai/provider/cloudflare

So they don't even support K3? What's the point. K2.7 Code is practically free already

They do:

Input tokens (per 1M)$3.00 Cached input tokens (per 1M)$0.30 Output tokens (per 1M)$15.00

Re: Smaller, faster, safer: running Kimi and GLM at scale

#40
post #26

> If squeezing the best open models onto GPUs and serving them to millions of developers sounds like your kind of problem, come work with us. What is the typical job title and/or skillset for this?

I've seen this called MLOps.

MLOps though typically wouldn't been going quantization? It requires some careful testing of accuracy and performance even today.

Applied ML Research Engineer or something maybe, not that I have ever seen that title. Maybe just catchall ML Engineer...

Post reply on HN