Live data from Hacker News

A new inference engine to run Kimi K3 2.78T parameter with 29GB of RAM

marcobambini.substack.com

1–3 of 3 posts

Re: A new inference engine to run Kimi K3 2.78T parameter with 29GB of RAM

#3

Current speed is “approximately one third of a token per second”

Right, we trade speed for the ability to run a 2.7T-parameter model while preserving accuracy. It is a first version, and we plan to improve the inference performance.