Live data from Hacker News

Self-hosting Kimi K3: 20% more hardware cost, 20% better task resolution

aistack.imec-int.com

1–10 of 56 posts

Re: Self-hosting Kimi K3: 20% more hardware cost, 20% better task resolution

#2
Author here(I'm on the team). We updated this post after Monday's Kimi K3 weights release: fitting the 2.8T model means going from 8×B200 to 8×B300, ~20% more hardware cost, and concurrency drops from 24 to 16 users vs GLM-5.2. Caveat we're upfront about in the post: our 64-task SWEBench Pro subset may be in Kimi's training set, so the 86% resolve rate is an upper bound.

Let us know your thoughts, we really value feedback!

Re: Self-hosting Kimi K3: 20% more hardware cost, 20% better task resolution

#3
>which box to buy

> spark, costs less than a conference trip.

I know putting actual prices regionally localizes your article and temporally, with how prices are so unstable. But analysis of “what to buy” without actual prices is borderline meaningless.

Overall, good article, very interesting to see a real deployment that’s actually attainable and not just a subscription to a big 3 token plan.

Re: Self-hosting Kimi K3: 20% more hardware cost, 20% better task resolution

#4
I would love to see such comparisons but with quantized versions, because quantization allows running models on smaller hardware with some quality loss. I am running Qwen3.6-35B-A3B quantized to int4 on an A6000 card that was otherwise just sitting around idle. It works up to a degree, but I would love to see benchmarks comparing different quantizations of several models, especially in quality. This is an important dimension in deciding whether to buy a GPU is worth it, and it is missing from this (otherwise comprehensive) article.

Re: Self-hosting Kimi K3: 20% more hardware cost, 20% better task resolution

#5
Recently I started playing with LM Studio and local models and found out that `gemma-4-26b-a4b` is suprisingly capable. I don't need elaborate akin to "create complete app to do X" or "refactor the whole codebase of bazzilion of LOC" but rather "how to go about doing thing X" or for language study (explaining nuances of phrasal verbs or subtelties of vocabulary in other languagues) and darn -- the results are rather good (to the point that I use it mostly nowadays)

Re: Self-hosting Kimi K3: 20% more hardware cost, 20% better task resolution

#6
This seems like a temporary situation. Utilization maximization is a matter of switching to the software factory unattended flow and intelligently routing tasks that will not require assistance to the unused time.

A decade and a half ago we used to run massive map reduce jobs overnight. Code will be handled like this.

Re: Self-hosting Kimi K3: 20% more hardware cost, 20% better task resolution

#7
I could not focus on the article with all the noise in the background. It was annoying in the header but then it continued down the page. If you want to do this on your marketing pages have at it but for a blog/news style page? Reader mode was the only way to restore sanity.

Re: Self-hosting Kimi K3: 20% more hardware cost, 20% better task resolution

#8

I could not focus on the article with all the noise in the background. It was annoying in the header but then it continued down the page. If you want to do this on your marketing pages have at it but for a blog/news style page? Reader mode was the only way to restore sanity.

Same here. Stopped reading.

Re: Self-hosting Kimi K3: 20% more hardware cost, 20% better task resolution

#9

I could not focus on the article with all the noise in the background. It was annoying in the header but then it continued down the page. If you want to do this on your marketing pages have at it but for a blog/news style page? Reader mode was the only way to restore sanity.

[flagged]

Re: Self-hosting Kimi K3: 20% more hardware cost, 20% better task resolution

#10
post #9

I could not focus on the article with all the noise in the background. It was annoying in the header but then it continued down the page. If you want to do this on your marketing pages have at it but for a blog/news style page? Reader mode was the only way to restore sanity.

[flagged]

> It should support reduced motion preference though for the more feeble among us.

What a strong start to a sentence before veering into a pretty gross equivalency.

Post reply on HN