Live data from Hacker News

Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model

moonshotai.github.io

431–440 of 442 posts

Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model

#431

As a Chinese user, I can say that many people use Kimi, even though I personally don’t use it much. China’s open-source strategy has many significant effects—not only because it aligns with the spirit of open source. For domestic Chinese companies, it also prevents startups from making reckless investments to develop mediocre models. Instead, everyone is pushed to start from a relatively high baseline. Of course, man…

How can I use it ?

google play or app store? or https://www.kimi.com/en/

Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model

#432
post #310

Earlier quoted context omitted.

Don't confuse the medium with the picture it represents. Porn is pornographic, whether it is a photo or an oil painting. Feelings are feelings, whether they're felt by a squishy meat brain or a perfect atom-by-atom simulation of one in a computer. Or a less-than-perfect simulation of one. Or just a vaguely similar system that is largely indistinguishable from it, as observed from the outside. Individual nerve cells d…

Do you think a simulation of a weather forcast is the same as the real weather? (And science fiction .. is not necessarily science)

> Do you think a simulation of a weather forcast is the same as the real weather?

If sufficiently accurate... then yes. It is weather.

We are mere information, encoded in the ripples of the fabric of the universe, nothing more.

Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model

#433
post #364

Earlier quoted context omitted.

you guys will outperform the US, no doubt. energy generation multiples of what the US is producing. What does AI need ? Energy. second - the open source nature of the models - means as you said a high baseline to start with - faster iteration.

> will outperform does outperform China is absolutely winning innovation in the 21st century. I'm so impressed. For an example from just this morning, there was an article that they're developing thorium reactor-powered cargo ships. I'm blown away.

Jm2c, but I really dislike those winners/losers narratives. They lack any nuance, are juvenile, and ultimately do not contribute much but noise like endless of pointless "who's better Jordan or Lebron?" debates.

Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model

#434

Earlier quoted context omitted.

I don't know if how close Europe is, but I'm sufficiently whelmed by Mistral that I don't need to look elsewhere yet. It's kind-of like having a Toyota Corolla while everybody else is driving around in smart cars but it gets it done. On top of it, there's a loyal community that (maybe because I'm not looking) I don't see with other products. It probably depends on your uses, but if I spent all my time chasing the lat…

> I don't know if how close Europe is, but I'm sufficiently whelmed by Mistral that I don't need to look elsewhere yet. It's kind-of like having a Toyota Corolla while everybody else is driving around in smart cars but it gets it done. My problem was that it really doesn't, none of the models out there are that great at agentic coding when you care about maintainability. Sonnet 4.5 sometimes struggles and is only oka…

Coding isn't the only use case.

Neither is being bleeding edge.

I use Mistral's models, I've built an entire internal-knowledge-pipeline of sort using Mistral's products (which involved anything from OCR, to summarization, to linking stuff across different services like Jira or Teams, etc) and I've been very happy with it.

We did consider alternatives and truth to be told none was as cost-effective, fast and satisfying (and also our company does not trust US AI companies to not do stuff with our data).

Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model

#435

Earlier quoted context omitted.

> I don't know if how close Europe is, but I'm sufficiently whelmed by Mistral that I don't need to look elsewhere yet. It's kind-of like having a Toyota Corolla while everybody else is driving around in smart cars but it gets it done. My problem was that it really doesn't, none of the models out there are that great at agentic coding when you care about maintainability. Sonnet 4.5 sometimes struggles and is only oka…

Coding isn't the only use case. Neither is being bleeding edge. I use Mistral's models, I've built an entire internal-knowledge-pipeline of sort using Mistral's products (which involved anything from OCR, to summarization, to linking stuff across different services like Jira or Teams, etc) and I've been very happy with it. We did consider alternatives and truth to be told none was as cost-effective, fast and satisfyi…

So you're not able to trust inference providers like Google Cloud w/ ZDR etc with your data?

Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model

#436
post #435

Earlier quoted context omitted.

Coding isn't the only use case. Neither is being bleeding edge. I use Mistral's models, I've built an entire internal-knowledge-pipeline of sort using Mistral's products (which involved anything from OCR, to summarization, to linking stuff across different services like Jira or Teams, etc) and I've been very happy with it. We did consider alternatives and truth to be told none was as cost-effective, fast and satisfyi…

So you're not able to trust inference providers like Google Cloud w/ ZDR etc with your data?

My EU-based clients are unwilling to do so as we see all clouds as black boxes you have no real idea what you getting into.

Most of our hosting is also on European providers, my team's the only one that deploys some services on Azure.

Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model

#437

Earlier quoted context omitted.

What 1T parameter base model have you seen from any of those labs?

its moe, each expert tower can be branched from some smaller model.

That's not how MoE works, you need to train the FFN directly or else the FFN gate would have no clue how to activate the expert.

Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model

#438

Earlier quoted context omitted.

> the whole business model of companies like OpenAI and Anthropic, at least at the moment, seems to be that the models are so big that you need to run them in the cloud with metered access. That's not a business model choice, though. That's a reality of running SOTA models. If OpenAI or Anthropic could squeeze the same output out of smaller GPUs and servers they'd be doing it for themselves. It would cut their datace…

> If OpenAI or Anthropic could squeeze the same output out of smaller GPUs and servers they'd be doing it for themselves. First, they do this; that's why they release models at different price points. It's also why GPT-5 tries auto-routing requests to the most cost-effective model. Second, be careful about considering the incentives of these companies. They all act as if they're in an existential race to deliver 'the…

> delivering 97% of the performance at 10% of the cost is a distraction.

Not if you are running RL on that model, and need to do many roll-outs.

Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model

#439

Earlier quoted context omitted.

I get a lot of meaning out of weights and source (without the training data), not sure about you. Calling it meaning less seems like exaggeration.

Can you change the weights to improve?

it's a bunch of numbers. Of course you can change them.

Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model

#440

Earlier quoted context omitted.

The electricity cost to run these models locally is already more than equivalent API cost.

That's going to depend on how small the model can be made, and how much you are using it. If we assume that running locally meant running on a 500W consumer GPU, then the electricity cost to run this non-stop 8 hours a day for 20 days a month (i.e. "business hours") would be around $10-20. This is about the same as OpenAI or Anthropics $20/mo plans, but for all day coding you would want their $100 or $200/mo plans, a…

Neither $20 nor $200 plans cover any API costs.

At $0.17 per million tokens for the smallest gpt model that's still faster rand more powerful than anything you can run locally and cheaper in kilowatts per hour than it would cost you to run locally even if you could.

Post reply on HN