Live data from Hacker News

Kimi Released Kimi K2.5, Open-Source Visual SOTA-Agentic Model

kimi.com

181–190 of 251 posts

Re: Kimi Released Kimi K2.5, Open-Source Visual SOTA-Agentic Model

#181

Earlier quoted context omitted.

Its often pointed out in the first sentence of a comment how a model can be run at home, then (maybe) towards the end of the comment it’s mentioned how it’s quantized. Back when 4k movies needed expensive hardware, no one was saying they could play 4k on a home system, then later mentioning they actually scaled down the resolution to make it possible. The degree of quality loss is not often characterized. Which makes…

The level of deceit you're describing is kind of ridiculous. Anybody talking about their specific setup is going to be happy to tell you the model and quant they're running and the speeds they're getting, and if you want to understand the effects of quantization on model quality, it's really easy to spin up a GPU server instance and play around.

> if you want to understand the effects of quantization on model quality, it's really easy to spin up a GPU server instance and play around

Fwiw, not necessarily. I've noticed quantized models have strange and surprising failure modes where everything seems to be working well and then does a death spiral repeating a specific word or completely failing on one task of a handful of similar tasks.

8-bit vs 4-bit can be almost imperceptible or night and day.

This isn't something you'd necessarily see playing around, but when trying to do something specific

Re: Kimi Released Kimi K2.5, Open-Source Visual SOTA-Agentic Model

#182
post #179

Earlier quoted context omitted.

We are getting into a debate between particulars and universals. To call the 'unified memory' VRAM is quite a generalization. Whatever the case, we can tell from stock prices that whatever this VRAM is, its nothing compared to NVIDIA. Anyway, we were trying to run a 70B model on a macbook(can't remember which M model) at a fortune 20 company, it never became practical. We were trying to compare strings of character l…

Here's a video of a previous 1T K2 model running using MLX on a a pair of Mac Studios: https://twitter.com/awnihannun/status/1943723599971443134 - performance isn't terrible.

Is there a catch? I was not getting anything like this on a 70B model.

EDIT: oh its a marketing account and the program never finished... who knows the validity.

Re: Kimi Released Kimi K2.5, Open-Source Visual SOTA-Agentic Model

#183
post #80
post #34

The "Deepseek moment" is just one year ago today! Coincidence or not, let's just marvel for a second over this amount of magic/technology that's being given away for free... and how liberating and different this is than OpenAI and others that were closed to "protect us all".

What amazes me is why would someone spend millions to train this model and give it away for free. What is the business here?

I think there is a book (Chip War) about how the USSR did not effectively participate in staying at the edge of the semiconductor revolution. And they have suffered for it.

China has decided they are going to participate in the LLM/AGI/etc revolution at any cost. So it is a sunk cost, and the models are just an end product and any revenue is validation and great, but not essential. The cheaper price points keep their models used and relevant. It challenges the other (US, EU) models to innovate and keep ahead to justify their higher valuations (both monthly plan, and investor). Once those advances are made, it can be bought back to their own models. In effect, the currently leading models are running from a second place candidate who never gets tired and eventually does what they do at a lower price point.

Re: Kimi Released Kimi K2.5, Open-Source Visual SOTA-Agentic Model

#184

Earlier quoted context omitted.

Hey have they open sourced all Kimi k2.5 (thinking,instruct,agent,agent swarm [beta])? Because I feel like they mentioned that agent swarm is available their api and that made me feel as if it wasn't open (weights)*? Please let me know if all are open source or not?

I'm assuming the swarm part is all harness. Well I mean a harness and way of thinking that the weights have just been fine tuned to use.

It's not in the harness today, it's a special RL technique they discuss in https://www.kimi.com/blog/kimi-k2-5.html (see "2. Agent Swarm")

I looked through the harness and all I could find is a `Task` tool.

Re: Kimi Released Kimi K2.5, Open-Source Visual SOTA-Agentic Model

#185
post #8

Huggingface Link: https://huggingface.co/moonshotai/Kimi-K2.5 1T parameters, 32b active parameters. License: MIT with the following modification: Our only modification part is that, if the Software (or any derivative works thereof) is used for any of your commercial products or services that have more than 100 million monthly active users, or more than 20 million US dollars (or equivalent in other currencies) in mont…

One. Trillion. Even on native int4 that’s… half a terabyte of vram?! Technical awe at this marvel aside that cracks the 50th percentile of HLE, the snarky part of me says there’s only half the danger in giving something away nobody can run at home anyway…

3,998.99 for 500gb of RAM on amazon

"Good Luck" - Kimi

Re: Kimi Released Kimi K2.5, Open-Source Visual SOTA-Agentic Model

#186
post #155

Earlier quoted context omitted.

The reason Macs get recommended is the unified memory, which is usable as VRAM for the GPU. People are similarly using the AMD Strix Halo for AI which also has a similar memory architecture. Time to first token for something like '1+1=' would be seconds, and then you'd be getting ~20 tokens per second, which is absolutely plenty fast for regular use. Token/s slows down at the higher end of context, but it's absolutel…

We are getting into a debate between particulars and universals. To call the 'unified memory' VRAM is quite a generalization. Whatever the case, we can tell from stock prices that whatever this VRAM is, its nothing compared to NVIDIA. Anyway, we were trying to run a 70B model on a macbook(can't remember which M model) at a fortune 20 company, it never became practical. We were trying to compare strings of character l…

With 32B active parameters, Kimi K2.5 will run faster than your 70B model.

Re: Kimi Released Kimi K2.5, Open-Source Visual SOTA-Agentic Model

#187
post #80

Earlier quoted context omitted.

What amazes me is why would someone spend millions to train this model and give it away for free. What is the business here?

I think there is a book (Chip War) about how the USSR did not effectively participate in staying at the edge of the semiconductor revolution. And they have suffered for it. China has decided they are going to participate in the LLM/AGI/etc revolution at any cost. So it is a sunk cost, and the models are just an end product and any revenue is validation and great, but not essential. The cheaper price points keep their…

In some way, the US won the cold war by spending so much on military that the USSR, in trying to keep up, collapsed. I don't see any parallels between that and China providing infinite free compute to their AI labs, why do you ask?

Re: Kimi Released Kimi K2.5, Open-Source Visual SOTA-Agentic Model

#188
post #8

Huggingface Link: https://huggingface.co/moonshotai/Kimi-K2.5 1T parameters, 32b active parameters. License: MIT with the following modification: Our only modification part is that, if the Software (or any derivative works thereof) is used for any of your commercial products or services that have more than 100 million monthly active users, or more than 20 million US dollars (or equivalent in other currencies) in mont…

Cursor devs, who go out of their way to not mention their Composer model is based on GLM, are not going to like that.

Source? I've heard this rumour twice but never seen proof. I assume it would be based on tokeniser quirks?

Re: Kimi Released Kimi K2.5, Open-Source Visual SOTA-Agentic Model

#189

Can we please stop calling those models "open source"? Yes the weights are open. So, "open weight" maybe. But the source isn't open, the thing that allows to re-create it. That's what "open source" used to mean. (Together with a license that allows you to use that source for various things.)

No major AI lab will admit to training on proprietary or copyrighted data so what you are asking is an impossibility. You can make a pretty good LLM if you train on Anna's Archive but it will either be released anonymously, or with a research only non commercial license.

There aren't enough public domain data to create good LLMs, especially once you get into the newer benchmarks that expect PhD level of domain expertise in various niche verticals.

It's also a logical impossibility to create a zero knowledge proof that will allow you to attribute to specific training data without admitting to usage.

I can think of a few technical options but none would hold water legally.

You can use a Σ-protocol OR-composition to prove that it was trained either on a copyrighted dataset or a non copyrighted dataset without admitting to which one (technically interesting, legally unsound).

You can prove that a model trained on copywrited data is statistically indistinguishable from one trained on non-copywrited data (an information theoretic impossibility unless there exist as much public domain data as copywrited data, in similar distributions).

You can prove a public domain and copywrited dataset are equivalent if the model performance produced is indistinguishable from each other.

All the proofs fail irl, ignoring the legal implications, because there's less public domain information, so given the lemma that more training data == improved model performance, all the above are close to impossible.

Re: Kimi Released Kimi K2.5, Open-Source Visual SOTA-Agentic Model

#190

Congratulations, great work Kimi team. Why is that Claude still at the top in coding, are they heavily focused on training for coding or is it their general training is so good that it performs well in coding? Someone please beat the Opus 4.5 in coding, I want to replace it.

Gemini 3 pro is way better than Opus especially for large codebases.

Do you use it only for code editing, or also for running bash commands? My experience is that it is very bad at the latter.
Post reply on HN