Live data from Hacker News

Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model

moonshotai.github.io

371–380 of 442 posts

Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model

#371

Earlier quoted context omitted.

Europe should act and make its own, literal, Moonshot: https://ifiwaspolitical.substack.com/p/euroai-europes-path-t...

>Moonshot 1: GPT-4 Parity (2027) >Objective: 100B parameter model matching GPT-4 benchmarks, proving European technical viability This feels like a joke... Parity with a 2024 model in 2027? The Chinese didn't wait, they just did it. The timeline for #1 LLM is also so far into the future that it is entirely plausible that by 2031, nobody uses transformer based LLMs as we know them today anymore. For reference: The att…

Note the EU-Moonshot project is based on own silicon / compute sovereignty.

GPT4 parity on a own silicon trained indigenous model is just an early goal.

Indeed, the ultimate goal is EU LLM supremacy - which means under democratic control.

Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model

#372
post #364

Earlier quoted context omitted.

you guys will outperform the US, no doubt. energy generation multiples of what the US is producing. What does AI need ? Energy. second - the open source nature of the models - means as you said a high baseline to start with - faster iteration.

> will outperform does outperform China is absolutely winning innovation in the 21st century. I'm so impressed. For an example from just this morning, there was an article that they're developing thorium reactor-powered cargo ships. I'm blown away.

I remember this thing. The tech is from America actually, decades ago. (Thorium). But they give up and china counties the work recent years

Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model

#373
post #220

Earlier quoted context omitted.

First it's an easy way to test censorship. Second, you might flip the question: why is the Chinese govt so obsessed that they still block all mention of the event?

I don’t get why the government doesn’t recognize the event and then mold it to its narrative, like so many other governments do. They basically need to give it the Hollywood treatment. I’m sure a lot of people don’t know that prior to the event, the protesters lynched and set soldiers on fire.

They do, but prefer to use their own keywords, such as the June 4th incident.

Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model

#374

Earlier quoted context omitted.

> will outperform does outperform China is absolutely winning innovation in the 21st century. I'm so impressed. For an example from just this morning, there was an article that they're developing thorium reactor-powered cargo ships. I'm blown away.

I remember this thing. The tech is from America actually, decades ago. (Thorium). But they give up and china counties the work recent years

"The tech is from America actually, decades ago... But they give up and china continues the work"

Many such cases...

Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model

#375
post #275

Earlier quoted context omitted.

> There’s honestly not a reason they have to be 1T parameters and cost an insane amount to train and run on inference. Kimi K2 Thinking is rumored to have cost $4.6m to train - according to "a source familiar with the matter": https://www.cnbc.com/2025/11/06/alibaba-backed-moonshot-rele... I think the most interesting recent Chinese model may be MiniMax M2, which is just 200B parameters but benchmarks close to Sonnet…

That number is as real as the 5.5 million to train DeepSeek. Maybe it's real if you're only counting the literal final training run, but total costs including the huge number of failed runs all other costs accounted for, it's several hundred million to train a model that's usually still worse than Claude, Gemini, or ChatGPT. It took 1B+ (500 billion on energy and chips ALONE) for Grok to get into the "big 4".

Using such theory, one can even argue that the real cost needs to include the infrastructures, like total investment into the semiconductor industry, the national electricity grid, education and even defence etc.

Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model

#376
post #364

Earlier quoted context omitted.

you guys will outperform the US, no doubt. energy generation multiples of what the US is producing. What does AI need ? Energy. second - the open source nature of the models - means as you said a high baseline to start with - faster iteration.

Going on a tangent, is Europe even close? Mistral has been underwhelming

I don't know if how close Europe is, but I'm sufficiently whelmed by Mistral that I don't need to look elsewhere yet. It's kind-of like having a Toyota Corolla while everybody else is driving around in smart cars but it gets it done. On top of it, there's a loyal community that (maybe because I'm not looking) I don't see with other products. It probably depends on your uses, but if I spent all my time chasing the latest chat models (like Kimi K2 for instance) I wouldn't actually get anything done.

Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model

#377

Earlier quoted context omitted.

> will outperform does outperform China is absolutely winning innovation in the 21st century. I'm so impressed. For an example from just this morning, there was an article that they're developing thorium reactor-powered cargo ships. I'm blown away.

I remember this thing. The tech is from America actually, decades ago. (Thorium). But they give up and china counties the work recent years

> The tech is from America actually, decades ago. (Thorium).

I guess it depends on how you see it, but regardless, the people putting it to use today doesn't seem to be in the US.

FWIW:

> Thorium was discovered in 1828 by the Swedish chemist Jöns Jacob Berzelius during his analysis of a new mineral [...] In 1824, after more deposits of the same mineral in Vest-Agder, Norway, were discovered [...] While thorium was discovered in 1828 its first application dates only from 1885, when Austrian chemist Carl Auer von Welsbach invented the gas mantle [...] Thorium was first observed to be radioactive in 1898, by the German chemist Gerhard Carl Schmidt

For being an American discovery, it sure has a lot of European people involved in it :) (I've said it elsewhere but worth repeating; trying to track down where a technology/invention actually comes from is a fools errand, and there is always something earlier that led to today, so doesn't serve much purpose except nationalism it seems to me).

Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model

#378
post #375

Earlier quoted context omitted.

That number is as real as the 5.5 million to train DeepSeek. Maybe it's real if you're only counting the literal final training run, but total costs including the huge number of failed runs all other costs accounted for, it's several hundred million to train a model that's usually still worse than Claude, Gemini, or ChatGPT. It took 1B+ (500 billion on energy and chips ALONE) for Grok to get into the "big 4".

Using such theory, one can even argue that the real cost needs to include the infrastructures, like total investment into the semiconductor industry, the national electricity grid, education and even defence etc.

Correct! You do have to account for all of these things! Unironically correct! :)

Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model

#379

Earlier quoted context omitted.

I get a lot of meaning out of weights and source (without the training data), not sure about you. Calling it meaning less seems like exaggeration.

Can you change the weights to improve?

You can fine tune without the original training data, which for a large LLM is typically going to mean using LoRA - keeping the original weights unchanged and adding separate fine-tuning weights.

Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model

#380
post #359

Earlier quoted context omitted.

I agree that they should say "open weight" instead of "open source" when that's what they mean, but it might take some time for people to understand that it's not the same thing exactly and we should allow some slack for that.

no. truly open source models are wonderful and remarkable things that truly move the needle in education, understanding, distributed collaboration and the advancement of the state of the art. redefinition of the terminology reduces incentive to strive for the wonderful goal that they represent.

There is a big difference between open source for something like the linux kernel or gcc where anyone with a home PC can build it, and any non-trivial LLM where it takes cloud compute and costs a lot to train it. No hobbyist or educational institution is going to be paying for million dollar training runs, probably not even thousand dollar ones.
Post reply on HN