Earlier quoted context omitted.
Its often pointed out in the first sentence of a comment how a model can be run at home, then (maybe) towards the end of the comment it’s mentioned how it’s quantized. Back when 4k movies needed expensive hardware, no one was saying they could play 4k on a home system, then later mentioning they actually scaled down the resolution to make it possible. The degree of quality loss is not often characterized. Which makes…
Didn't this paper demonstrate that you only need 1.58 bits to be equivalent to 16 bits in performance? https://arxiv.org/abs/2402.17764
Kimi Released Kimi K2.5, Open-Source Visual SOTA-Agentic Model
221–230 of 251 posts
Re: Kimi Released Kimi K2.5, Open-Source Visual SOTA-Agentic Model
#222Huggingface Link: https://huggingface.co/moonshotai/Kimi-K2.5 1T parameters, 32b active parameters. License: MIT with the following modification: Our only modification part is that, if the Software (or any derivative works thereof) is used for any of your commercial products or services that have more than 100 million monthly active users, or more than 20 million US dollars (or equivalent in other currencies) in mont…
One. Trillion. Even on native int4 that’s… half a terabyte of vram?! Technical awe at this marvel aside that cracks the 50th percentile of HLE, the snarky part of me says there’s only half the danger in giving something away nobody can run at home anyway…
Kimi-K2-Thinking-UD-Q3_K_XL-00001-of-00010.gguf Generation - 5,231 tokens 604.63s 8.65 tokens/s
Re: Kimi Released Kimi K2.5, Open-Source Visual SOTA-Agentic Model
#223The "Deepseek moment" is just one year ago today! Coincidence or not, let's just marvel for a second over this amount of magic/technology that's being given away for free... and how liberating and different this is than OpenAI and others that were closed to "protect us all".
Re: Kimi Released Kimi K2.5, Open-Source Visual SOTA-Agentic Model
#224Is this actually good or just optimized heavily for benchmarks? I am hopefully its the former based on the writeup but need to put it through its paces.
Re: Kimi Released Kimi K2.5, Open-Source Visual SOTA-Agentic Model
#225Earlier quoted context omitted.
Hear, hear. Even if the model fits, a few tokens per second make no sense. Time is money too.
If I can start an agent and be able to walk away for 8 hours, and be confident it's 'smart' enough to complete a task unattended, that's still useful. At 3 tk/s, that's still 100-150 pages of a book, give or take.
Re: Kimi Released Kimi K2.5, Open-Source Visual SOTA-Agentic Model
#226Earlier quoted context omitted.
One. Trillion. Even on native int4 that’s… half a terabyte of vram?! Technical awe at this marvel aside that cracks the 50th percentile of HLE, the snarky part of me says there’s only half the danger in giving something away nobody can run at home anyway…
I run KimiK2 at home, Most of it on system ram with a few layers offloaded to old 3090s. This is a cheap budget build. Kimi-K2-Thinking-UD-Q3_K_XL-00001-of-00010.gguf Generation - 5,231 tokens 604.63s 8.65 tokens/s
I currently have a 3970x with a bunch of 3090s.
Re: Kimi Released Kimi K2.5, Open-Source Visual SOTA-Agentic Model
#227Earlier quoted context omitted.
I don't think Awni should be dismissed as a "marketing account" - they're an engineer at Apple who's been driving the MLX project for a couple of years now, they've earned a lot of respect from me.
Given how secretive Apple is, oh my, its super duper marketing account.
The tooling involved has improved significantly over the past year.
Re: Kimi Released Kimi K2.5, Open-Source Visual SOTA-Agentic Model
#228Earlier quoted context omitted.
Local LLMs are just LLMs people run locally. It's not a definition of size, feature set, or what's most popular. What the "real" value is for local LLMs will depend on each person you ask. The person who runs small local LLMs will tell you the real value is in small models, the person who runs large local LLMs will tell you it's large ones, those who use cloud will say the value is in shared compute, and those who do…
> LLMs which the weights aren't available are an example of when it's not local LLMs, not when the model happens to be large. I agree. My point was that most aren't thinking of models this large when they're talking about local LLMs. That's what I said, right? This is supported by the download counts on hf: the most downloaded local models are significantly smaller than 1tln, normally 1 - 12bln. I'm not sure I unders…
Re: Kimi Released Kimi K2.5, Open-Source Visual SOTA-Agentic Model
#229Earlier quoted context omitted.
Local LLMs are just LLMs people run locally. It's not a definition of size, feature set, or what's most popular. What the "real" value is for local LLMs will depend on each person you ask. The person who runs small local LLMs will tell you the real value is in small models, the person who runs large local LLMs will tell you it's large ones, those who use cloud will say the value is in shared compute, and those who do…
> LLMs which the weights aren't available are an example of when it's not local LLMs, not when the model happens to be large. I agree. My point was that most aren't thinking of models this large when they're talking about local LLMs. That's what I said, right? This is supported by the download counts on hf: the most downloaded local models are significantly smaller than 1tln, normally 1 - 12bln. I'm not sure I unders…
Re: Kimi Released Kimi K2.5, Open-Source Visual SOTA-Agentic Model
#230Earlier quoted context omitted.
Local LLMs are just LLMs people run locally. It's not a definition of size, feature set, or what's most popular. What the "real" value is for local LLMs will depend on each person you ask. The person who runs small local LLMs will tell you the real value is in small models, the person who runs large local LLMs will tell you it's large ones, those who use cloud will say the value is in shared compute, and those who do…
> LLMs which the weights aren't available are an example of when it's not local LLMs, not when the model happens to be large. I agree. My point was that most aren't thinking of models this large when they're talking about local LLMs. That's what I said, right? This is supported by the download counts on hf: the most downloaded local models are significantly smaller than 1tln, normally 1 - 12bln. I'm not sure I unders…
You're of course accurate that smaller LLMs are more commonly deployed, it's just not the part I was really responding to.