While I absolutely support these open source models, there is an interesting angle to consider... If I were a Chinese partisan looking to inflict a devastating blow to the US, taking the AI hype wind out of American tech valuation sails would seem a great option. How best to do this? Release highly performant models... For free! Extremely efficient in terms of RMB spent vs (unrealized) USD lost. But surely, these mod…
But they’re literally not free. If it was “war”, with infinite money to throw at destruction of USA AI industry, then why would you be charging and reducing such an outcome
Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model
271–280 of 442 posts
Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model
#272As a Chinese user, I can say that many people use Kimi, even though I personally don’t use it much. China’s open-source strategy has many significant effects—not only because it aligns with the spirit of open source. For domestic Chinese companies, it also prevents startups from making reckless investments to develop mediocre models. Instead, everyone is pushed to start from a relatively high baseline. Of course, man…
There’s a lot of indications that we’re currently brute forcing these models. There’s honestly not a reason they have to be 1T parameters and cost an insane amount to train and run on inference. What we’re going to see is as energy becomes a problem; they’ll simply shift to more effective and efficient architectures on both physical hardware and model design. I suspect they can also simply charge more for the service…
This is much more likely to be an issue in the US than in China. https://fortune.com/2025/08/14/data-centers-china-grid-us-in...
Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model
#273While I absolutely support these open source models, there is an interesting angle to consider... If I were a Chinese partisan looking to inflict a devastating blow to the US, taking the AI hype wind out of American tech valuation sails would seem a great option. How best to do this? Release highly performant models... For free! Extremely efficient in terms of RMB spent vs (unrealized) USD lost. But surely, these mod…
Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model
#274Earlier quoted context omitted.
There’s a lot of indications that we’re currently brute forcing these models. There’s honestly not a reason they have to be 1T parameters and cost an insane amount to train and run on inference. What we’re going to see is as energy becomes a problem; they’ll simply shift to more effective and efficient architectures on both physical hardware and model design. I suspect they can also simply charge more for the service…
> What we’re going to see is as energy becomes a problem This is much more likely to be an issue in the US than in China. https://fortune.com/2025/08/14/data-centers-china-grid-us-in...
https://www.brookings.edu/articles/how-do-china-and-america-...
The source for China's energy is more fragile than that of the US.
> Coal is by far China’s largest energy source, while the United States has a more balanced energy system, running on roughly one-third oil, one-third natural gas, and one-third other sources, including coal, nuclear, hydroelectricity, and other renewables.
Also, China's GDP is a bit less inefficient in terms of power used per unit of GDP. China relies on coal and imports.
> However, China uses roughly 20% more energy per unit of GDP than the United States.
Remember, China still suffers from blackouts due to manufacturing demand not matching supply. The fortune article seems like a fluff piece.
https://www.npr.org/2021/10/01/1042209223/why-covid-is-affec...
Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model
#275As a Chinese user, I can say that many people use Kimi, even though I personally don’t use it much. China’s open-source strategy has many significant effects—not only because it aligns with the spirit of open source. For domestic Chinese companies, it also prevents startups from making reckless investments to develop mediocre models. Instead, everyone is pushed to start from a relatively high baseline. Of course, man…
There’s a lot of indications that we’re currently brute forcing these models. There’s honestly not a reason they have to be 1T parameters and cost an insane amount to train and run on inference. What we’re going to see is as energy becomes a problem; they’ll simply shift to more effective and efficient architectures on both physical hardware and model design. I suspect they can also simply charge more for the service…
Kimi K2 Thinking is rumored to have cost $4.6m to train - according to "a source familiar with the matter": https://www.cnbc.com/2025/11/06/alibaba-backed-moonshot-rele...
I think the most interesting recent Chinese model may be MiniMax M2, which is just 200B parameters but benchmarks close to Sonnet 4, at least for coding. That's small enough to run well on ~$5,000 of hardware, as opposed to the 1T models which require vastly more expensive machines.
Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model
#276Earlier quoted context omitted.
I must be missing something important here. How do the Chinese train these models if they don't have access to the GPUs to train them?
> How do the Chinese train these models if they don't have access to the GPUs to train them? they may be taking some western models: llama, chatgpt-oss, gemma, mistral, etc, and do postraining, which required way less resources.
Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model
#277Earlier quoted context omitted.
> How do the Chinese train these models if they don't have access to the GPUs to train them? they may be taking some western models: llama, chatgpt-oss, gemma, mistral, etc, and do postraining, which required way less resources.
If they were doing that I expect someone would have found evidence of it. Everything I've seen so far has lead me to believe that these Chinese AI labs are training their own models from scratch.
Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model
#278How does one effectively use something like this locally with consumer-grade hardware?
Update: https://huggingface.co/mlx-community/Kimi-K2-Thinking - and here it is running on two M3 Ultras: https://x.com/awnihannun/status/1986601104130646266
Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model
#279Earlier quoted context omitted.
i disagree. words matter. the whole point of open source is that anyone can look and see exactly how the sausage is made. that is the point. that is why the word "open" is used. ...and sure, compiling gcc is nondeterministic too, but i can still inspect the complete source from where it comes because it is open source, which means that all of the source materials are available for inspection.
The point of open source in software is as you say. It's just not the same thing though. Using words and phrases differently in different fields is common.
the practice of science itself would be far stronger if it took more pages from open source software culture.
Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model
#280Earlier quoted context omitted.
> What we’re going to see is as energy becomes a problem This is much more likely to be an issue in the US than in China. https://fortune.com/2025/08/14/data-centers-china-grid-us-in...
Disagree. Part of the reason China produces more power (and pollution) is due to China manufacturing for the US. https://www.brookings.edu/articles/how-do-china-and-america-... The source for China's energy is more fragile than that of the US. > Coal is by far China’s largest energy source, while the United States has a more balanced energy system, running on roughly one-third oil, one-third natural gas, and one-thir…
China has been adding something like a 1GW coal plant’s worth of solar generation every eight hours in the past year, and the rate is accelerating. The US is no longer a serious competitor for China when it comes to energy production.