Live data from Hacker News

Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model

moonshotai.github.io

411–420 of 442 posts

Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model

#412

Earlier quoted context omitted.

Going on a tangent, is Europe even close? Mistral has been underwhelming

I don't know if how close Europe is, but I'm sufficiently whelmed by Mistral that I don't need to look elsewhere yet. It's kind-of like having a Toyota Corolla while everybody else is driving around in smart cars but it gets it done. On top of it, there's a loyal community that (maybe because I'm not looking) I don't see with other products. It probably depends on your uses, but if I spent all my time chasing the lat…

You have to try the latest Corolla then. Really smart. Lane and collision assistance, ... Unlike my old Corolla which is total dumb. It even doesn't turn the light off when I leave the car

Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model

#413
post #209
post #33

Earlier quoted context omitted.

"open source" means there should be a script that downloads all the training materials and then spins up a pipeline that trains end to end. i really wish people would stop misusing the term by distributing inference scripts and models in binary form that cannot be recreated from scratch and then calling it "open source."

The meaning of Open Source 1990: Free Software 2000: Open Source: Finally we sanitized ourselves of that activism! It was scaring away customers! 2010: Source is available (under our very restrictive license) 2020: What source?

2025: What prompt?

Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model

#414
post #332

Earlier quoted context omitted.

> China is the only country that is not controlled by ideology and is increasing its electricity capacity in a scientific way. Have recently noticed a lot of pro-CCP propaganda on social media (especially Instagram and TikTok), but strangely also on HN; kind of interesting. To anyone making the (trivially false) claim that China is not controlled by ideology, I'm not quite sure how you'd convince them of the opposite…

I mean only on this specific topic: electricity. Arguing with other things is pointless since HN has the same political leaning as reddit so I will pass

What's your Reddit username? I'm interested in reading your posts there.

Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model

#415

Can't wait for Artificial analysis benchmarks, still waiting on them adding Qwen3-max thinking, will be interesting to see how these two compare to each other

The analysis is up! Impressive: https://artificialanalysis.ai/models/kimi-k2-thinking

Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model

#416

Earlier quoted context omitted.

I mean only on this specific topic: electricity. Arguing with other things is pointless since HN has the same political leaning as reddit so I will pass

What's your Reddit username? I'm interested in reading your posts there.

I don’t have one now. I used to post lots of comments on china stuff but I got banned once and every time I registered a new one it will be banned soon. I guess they banned all my ip. So I only go anonymous now

Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model

#417
post #415

Can't wait for Artificial analysis benchmarks, still waiting on them adding Qwen3-max thinking, will be interesting to see how these two compare to each other

The analysis is up! Impressive: https://artificialanalysis.ai/models/kimi-k2-thinking

Wow, these numbers are insanse! I tried it yesterday and it worked beautifully well. It also responded the way I wanted every time, I didn't have to spend time prompting it on how to respond properly (unlike Grok 4 expert, which tends to yap a lot), it just knew.

Todays models have gotten so good that at this point, whatever I run, just works and helps me in whatever. Maybe I should start noting down prompts that some models fails at.

Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model

#420

Earlier quoted context omitted.

Going on a tangent, is Europe even close? Mistral has been underwhelming

I don't know if how close Europe is, but I'm sufficiently whelmed by Mistral that I don't need to look elsewhere yet. It's kind-of like having a Toyota Corolla while everybody else is driving around in smart cars but it gets it done. On top of it, there's a loyal community that (maybe because I'm not looking) I don't see with other products. It probably depends on your uses, but if I spent all my time chasing the lat…

> I don't know if how close Europe is, but I'm sufficiently whelmed by Mistral that I don't need to look elsewhere yet. It's kind-of like having a Toyota Corolla while everybody else is driving around in smart cars but it gets it done.

My problem was that it really doesn't, none of the models out there are that great at agentic coding when you care about maintainability. Sonnet 4.5 sometimes struggles and is only okay with some steering, same for Gemini Pro 2.5, GPT-5 recently seems like it's closer to "just working" with high reasoning, but still is expensive and slow. Cerebras recently started offering GLM-4.6 and it's roughly on par with Sonnet 4 so not great, but 24M tokens per day for 50 USD seems like good value even with 128k context limitation.

I don't think there is a single model that is good enough and dependable enough in my experience out there yet, I'll probably keep jumping around for the next 5-10 years (assuming the models keep improving until we hit diminishing returns so hard that it all evens out, hopefully after they've reached a satisfying baseline usefulness).

Don't get me wrong, all of those models can already provide value, it's just that they're pretty finnicky a lot of the time, some of it inherent due to how LLMs work, but some of it also because they should just be trained better and more. And the tools they're given should be better. And the context should be managed better. And I shouldn't see something as simple as diffs fail to apply repeatedly just because I'm asking for 100% accuracy in the search/replace to avoid them messing up the brackets or whatever else.

Post reply on HN