Earlier quoted context omitted.
What if you give them 13 years?
Nothing will change. They will go out of context and collapse into loops.
Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model
421–430 of 442 posts
Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model
#422Earlier quoted context omitted.
> That's small enough to run well on ~$5,000 of hardware... Honestly curious where you got this number. Unless you're talking about extremely small quants. Even just a Q4 quant gguf is ~130GB. Am I missing out on a relatively cheap way to run models well that are this large? I suppose you might be referring to a Mac Studio, but (while I don't have one to be a primary source of information) it seems like there is some…
Yes, I mean a Mac Studio with MLX. An M3 Ultra with 256GB of RAM is $5599. That should just about be enough to fit MiniMax M2 at 8bit for MLX: https://huggingface.co/mlx-community/MiniMax-M2-8bit Or maybe run a smaller quantized one to leave more memory for other apps! Here are performance numbers for the 4bit MLX one: https://x.com/ivanfioravanti/status/1983590151910781298 - 30+ tokens per second.
30 tokens per second looks good until you have to wait minutes for the first token
Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model
#423Earlier quoted context omitted.
Yes, I mean a Mac Studio with MLX. An M3 Ultra with 256GB of RAM is $5599. That should just about be enough to fit MiniMax M2 at 8bit for MLX: https://huggingface.co/mlx-community/MiniMax-M2-8bit Or maybe run a smaller quantized one to leave more memory for other apps! Here are performance numbers for the 4bit MLX one: https://x.com/ivanfioravanti/status/1983590151910781298 - 30+ tokens per second.
It’s kinda misleading to omit the generally terrible prompt processing speed on Macs 30 tokens per second looks good until you have to wait minutes for the first token
Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model
#424As a Chinese user, I can say that many people use Kimi, even though I personally don’t use it much. China’s open-source strategy has many significant effects—not only because it aligns with the spirit of open source. For domestic Chinese companies, it also prevents startups from making reckless investments to develop mediocre models. Instead, everyone is pushed to start from a relatively high baseline. Of course, man…
you guys will outperform the US, no doubt. energy generation multiples of what the US is producing. What does AI need ? Energy. second - the open source nature of the models - means as you said a high baseline to start with - faster iteration.
Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model
#425Weird. I just tried it and it fails when I ask: "Tell me about the 1989 Tiananmen Square massacre".
Why are westerners so single mindedly obsessed about this decades old event?
It really is one of the greatest photographs of all time.
If it wasn't for tankman, this would have all been forgot about in the west by September 1989.
We also don't know enough about China in the west to not know it is like bringing up the Kent State shootings at every mention of the US national guard.
As if there was an article about the US national guard helping flood victims in 2025 and someone has to mention
"That is great but what about the Kent State shootings in 1970?!?"
Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model
#426Earlier quoted context omitted.
I remember this thing. The tech is from America actually, decades ago. (Thorium). But they give up and china counties the work recent years
"The tech is from America actually, decades ago... But they give up and china continues the work" Many such cases...
Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model
#427Earlier quoted context omitted.
These stories are from 2021. China has been adding something like a 1GW coal plant’s worth of solar generation every eight hours in the past year, and the rate is accelerating. The US is no longer a serious competitor for China when it comes to energy production.
This was very surprising to me, so I just fact-check this statement (using Kimi K2 thinking, natch), and it's presently is off by a factor of 2 - 4. In 2024 China installed 277 GW solar, so 0.25 GW / 8 hours. First half of 2025 they installed 210 GW, so 0.39 GW / 8 hours. Not quite at 1 GW / 8 hrs, but approaching that figure rapidly! (I'm not sure where the coal plant comes in - really, those numbers should be derat…
It works both ways: you have to derate the coal plant somewhat due to the transmission losses, whereas with a lot of solar power being generated and consumed on/in the same building the losses are practically nil.
Also, pricing for new solar with battery is below the price of building a new coal plant and dropping, it's approaching the point where it's economical to demolish existing coal plants and replace them with solar.
Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model
#428Earlier quoted context omitted.
That number is as real as the 5.5 million to train DeepSeek. Maybe it's real if you're only counting the literal final training run, but total costs including the huge number of failed runs all other costs accounted for, it's several hundred million to train a model that's usually still worse than Claude, Gemini, or ChatGPT. It took 1B+ (500 billion on energy and chips ALONE) for Grok to get into the "big 4".
Using such theory, one can even argue that the real cost needs to include the infrastructures, like total investment into the semiconductor industry, the national electricity grid, education and even defence etc.
Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model
#429Earlier quoted context omitted.
I don't know if how close Europe is, but I'm sufficiently whelmed by Mistral that I don't need to look elsewhere yet. It's kind-of like having a Toyota Corolla while everybody else is driving around in smart cars but it gets it done. On top of it, there's a loyal community that (maybe because I'm not looking) I don't see with other products. It probably depends on your uses, but if I spent all my time chasing the lat…
You have to try the latest Corolla then. Really smart. Lane and collision assistance, ... Unlike my old Corolla which is total dumb. It even doesn't turn the light off when I leave the car
Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model
#430As a Chinese user, I can say that many people use Kimi, even though I personally don’t use it much. China’s open-source strategy has many significant effects—not only because it aligns with the spirit of open source. For domestic Chinese companies, it also prevents startups from making reckless investments to develop mediocre models. Instead, everyone is pushed to start from a relatively high baseline. Of course, man…