Live data from Hacker News

Self-hosting Kimi K3: 20% more hardware cost, 20% better task resolution

aistack.imec-int.com

51–56 of 56 posts

Re: Self-hosting Kimi K3: 20% more hardware cost, 20% better task resolution

#51

Earlier quoted context omitted.

If the memory price don't go down anymore, we could say that we are already priced out.

The memory prices put computers into the same category as cars. Expensive but still realistic. It's these 200kUSD+ nodes with a dozen or more data center class GPUs that are killing the self-hosting dream. The cheapest nodes are is in the same price category as literal houses. Either we get some competition in this space, or computing will go back to their roots as ivory tower big iron.

HBF will probably save the small inference nodes. (Eventually, the initial price will obviously be in the stratosphere.)

Long term, HBF shouldn't be more than 4x as expensive as commodity flash, and Kimi K3 is an interesting target for a system using it given how aggressively it compresses the KV-cache. An inference box with ~1.5TB of HBF and ~48GB of DRAM should be able to be built for less than a couple of grand, and get something near to 100tok/s on full Kimi K3 for a single token stream.

Re: Self-hosting Kimi K3: 20% more hardware cost, 20% better task resolution

#52

Earlier quoted context omitted.

You can get easily over 100 tok/s on the Gemma 4 31B QAT if you enable MTP. Same goes for the 27B. I'm getting 880 tok/s of throughout on a single RTX 5090 for batch tasks.

I've switched over to unsloth studio and am hitting those speeds now, thanks for the tip.

Unsloth studio is just llama.cpp/llama-server under the hood, so you should see the same performance with the latest daily compiled llama-server (and the right CLI options to load a GGUF file) and any of your own choice of tooling on top of it.

Re: Self-hosting Kimi K3: 20% more hardware cost, 20% better task resolution

#54

I would love to see such comparisons but with quantized versions, because quantization allows running models on smaller hardware with some quality loss. I am running Qwen3.6-35B-A3B quantized to int4 on an A6000 card that was otherwise just sitting around idle. It works up to a degree, but I would love to see benchmarks comparing different quantizations of several models, especially in quality. This is an important d…

the a3b does not need to be all in your card. it being a moe you could trivially run it q8 / fp8 if you have the ram

Re: Self-hosting Kimi K3: 20% more hardware cost, 20% better task resolution

#55
post #36

Earlier quoted context omitted.

Don't worry, you'll be able to rent your next computer: https://www.apple.com/newsroom/2026/07/apple-upgrade-launche...

It's not like people ever owned iPhones to begin with. They're Apple's computers, Apple's just generously allowing their customers to use them, and only on their terms. A monthly iPhone subscription is just the end game. You will nothing, and you'll be happy.

Sure, but it's an actual computer I am holding and can use, whereas we are soon to have cloud personal computing.
Post reply on HN