Live data from Hacker News

Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model

moonshotai.github.io

401–410 of 442 posts

Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model

#401
post #275

Earlier quoted context omitted.

There’s a lot of indications that we’re currently brute forcing these models. There’s honestly not a reason they have to be 1T parameters and cost an insane amount to train and run on inference. What we’re going to see is as energy becomes a problem; they’ll simply shift to more effective and efficient architectures on both physical hardware and model design. I suspect they can also simply charge more for the service…

> There’s honestly not a reason they have to be 1T parameters and cost an insane amount to train and run on inference. Kimi K2 Thinking is rumored to have cost $4.6m to train - according to "a source familiar with the matter": https://www.cnbc.com/2025/11/06/alibaba-backed-moonshot-rele... I think the most interesting recent Chinese model may be MiniMax M2, which is just 200B parameters but benchmarks close to Sonnet…

> That's small enough to run well on ~$5,000 of hardware...

Honestly curious where you got this number. Unless you're talking about extremely small quants. Even just a Q4 quant gguf is ~130GB. Am I missing out on a relatively cheap way to run models well that are this large?

I suppose you might be referring to a Mac Studio, but (while I don't have one to be a primary source of information) it seems like there is some argument to be made on whether they run models "well"?

Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model

#402
post #275

Earlier quoted context omitted.

> There’s honestly not a reason they have to be 1T parameters and cost an insane amount to train and run on inference. Kimi K2 Thinking is rumored to have cost $4.6m to train - according to "a source familiar with the matter": https://www.cnbc.com/2025/11/06/alibaba-backed-moonshot-rele... I think the most interesting recent Chinese model may be MiniMax M2, which is just 200B parameters but benchmarks close to Sonnet…

> That's small enough to run well on ~$5,000 of hardware... Honestly curious where you got this number. Unless you're talking about extremely small quants. Even just a Q4 quant gguf is ~130GB. Am I missing out on a relatively cheap way to run models well that are this large? I suppose you might be referring to a Mac Studio, but (while I don't have one to be a primary source of information) it seems like there is some…

Yes, I mean a Mac Studio with MLX.

An M3 Ultra with 256GB of RAM is $5599. That should just about be enough to fit MiniMax M2 at 8bit for MLX: https://huggingface.co/mlx-community/MiniMax-M2-8bit

Or maybe run a smaller quantized one to leave more memory for other apps!

Here are performance numbers for the 4bit MLX one: https://x.com/ivanfioravanti/status/1983590151910781298 - 30+ tokens per second.

Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model

#403
post #382

Earlier quoted context omitted.

That’s backwards. New research and ideas are proven on small models. Lots and lots of ideas are tested that way. Good ideas get scaled up to show they still work on medium sized models. The very best ideas make their way into the code for the next huge training runs, which can cost tens or hundreds of millions of dollars. Not to nitpick words, but ablation is the practice of stripping out features of an algorithm or…

> Not to nitpick words, but ablation is the practice of stripping out features of an algorithm ... Ablation generally refers to removing parts of a system to see how it performs without them. In the context of an LLM it can refer to training data as well as the model itself. I'm not saying it'd be the most cost-effective method, but one could certainly try to create a small coding model by starting with a large one t…

ML researchers will sometimes vary the size of the training data set to see what happens. It’s not common - except in scaling law research. But it’s never called “ablation”.

Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model

#404

Earlier quoted context omitted.

"I don't understand. We already have that capability in our skulls. It's also "already there", so it would be a waste to not use it." seems like you are here that not understand this Company want to replace human and won't need to pay massive salary

I understand the companies wanting it. I hate it, but I understand. I don’t understand the humans wanting to be replaced though.

"I don’t understand the humans wanting to be replaced though."

because human that replace these job isnt the same human that got cut????

human that can replace these jobs would be rich

Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model

#405
post #275

Earlier quoted context omitted.

> There’s honestly not a reason they have to be 1T parameters and cost an insane amount to train and run on inference. Kimi K2 Thinking is rumored to have cost $4.6m to train - according to "a source familiar with the matter": https://www.cnbc.com/2025/11/06/alibaba-backed-moonshot-rele... I think the most interesting recent Chinese model may be MiniMax M2, which is just 200B parameters but benchmarks close to Sonnet…

> That's small enough to run well on ~$5,000 of hardware... Honestly curious where you got this number. Unless you're talking about extremely small quants. Even just a Q4 quant gguf is ~130GB. Am I missing out on a relatively cheap way to run models well that are this large? I suppose you might be referring to a Mac Studio, but (while I don't have one to be a primary source of information) it seems like there is some…

Running in cpu ram works fine. It’s not hard to build a machine with a terabyte of RAM.

Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model

#406

As a Chinese user, I can say that many people use Kimi, even though I personally don’t use it much. China’s open-source strategy has many significant effects—not only because it aligns with the spirit of open source. For domestic Chinese companies, it also prevents startups from making reckless investments to develop mediocre models. Instead, everyone is pushed to start from a relatively high baseline. Of course, man…

[dead]

Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model

#407
post #286

Earlier quoted context omitted.

Disagree. Part of the reason China produces more power (and pollution) is due to China manufacturing for the US. https://www.brookings.edu/articles/how-do-china-and-america-... The source for China's energy is more fragile than that of the US. > Coal is by far China’s largest energy source, while the United States has a more balanced energy system, running on roughly one-third oil, one-third natural gas, and one-thir…

China’s breakneck development is difficult for many in the US to grasp (root causes - baselining on sluggish domestic growth, and possessing a condescending view of China). This article offers a far more accurate picture than of how China is doing right now: https://archive.is/wZes6

Thank you for sharing this article. Eye opening.

Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model

#408

Earlier quoted context omitted.

Disagree. Part of the reason China produces more power (and pollution) is due to China manufacturing for the US. https://www.brookings.edu/articles/how-do-china-and-america-... The source for China's energy is more fragile than that of the US. > Coal is by far China’s largest energy source, while the United States has a more balanced energy system, running on roughly one-third oil, one-third natural gas, and one-thir…

These stories are from 2021. China has been adding something like a 1GW coal plant’s worth of solar generation every eight hours in the past year, and the rate is accelerating. The US is no longer a serious competitor for China when it comes to energy production.

This was very surprising to me, so I just fact-check this statement (using Kimi K2 thinking, natch), and it's presently is off by a factor of 2 - 4. In 2024 China installed 277 GW solar, so 0.25 GW / 8 hours. First half of 2025 they installed 210 GW, so 0.39 GW / 8 hours.

Not quite at 1 GW / 8 hrs, but approaching that figure rapidly!

(I'm not sure where the coal plant comes in - really, those numbers should be derated relative to a coal plant, which can run 24/7)

Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model

#409
post #405

Earlier quoted context omitted.

> That's small enough to run well on ~$5,000 of hardware... Honestly curious where you got this number. Unless you're talking about extremely small quants. Even just a Q4 quant gguf is ~130GB. Am I missing out on a relatively cheap way to run models well that are this large? I suppose you might be referring to a Mac Studio, but (while I don't have one to be a primary source of information) it seems like there is some…

Running in cpu ram works fine. It’s not hard to build a machine with a terabyte of RAM.

Admittedly I've not tried running on system RAM often, but every time I've tried it's been abysmally slow (< 1 T/s) when I've tried on something like KoboldCPP or ollama. Is there any particular method required to run them faster? Or is it just "get faster RAM"? I fully admit my DDR3 system has quite slow RAM...

Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model

#410
post #402

Earlier quoted context omitted.

> That's small enough to run well on ~$5,000 of hardware... Honestly curious where you got this number. Unless you're talking about extremely small quants. Even just a Q4 quant gguf is ~130GB. Am I missing out on a relatively cheap way to run models well that are this large? I suppose you might be referring to a Mac Studio, but (while I don't have one to be a primary source of information) it seems like there is some…

Yes, I mean a Mac Studio with MLX. An M3 Ultra with 256GB of RAM is $5599. That should just about be enough to fit MiniMax M2 at 8bit for MLX: https://huggingface.co/mlx-community/MiniMax-M2-8bit Or maybe run a smaller quantized one to leave more memory for other apps! Here are performance numbers for the 4bit MLX one: https://x.com/ivanfioravanti/status/1983590151910781298 - 30+ tokens per second.

Thanks for the info! Definitely much better than I expected.
Post reply on HN