Live data from Hacker News

Kimi-K3 on HuggingFace

huggingface.co

371–380 of 588 posts

Re: Kimi-K3 on HuggingFace

#371
post #129

Earlier quoted context omitted.

You're right, and it's interesting to consider why. It's probably a combination of a few factors: 1) Local LLMs are a relatively new phenomenon and hardware takes years. Apple probably lucked into their unified memory architecture being suitable (in terms of memory size and bandwidth) for local LLMs, but it's only with the newest generations we're hearing about LLMs even being a consideration in their design process.…

There will be a huge market for local inference once it's cheap and widely available. Try to imagine output token speeds of 15,000 tok/s and a time-to-first-token of 200ms. (This has already been done for Llama 8B.) Now imagine gargantuan context windows (2M, 4M, or even bigger); keep in mind the 1M context windows were science fiction a few years ago... now imagine having this on a local model on something like a ph…

> There will be a huge market for local inference once it's cheap and widely available.

I've seen public pronouncements that the RAM shortage could persist for a decade.

And then if consider that the constraint on local LLMs isn't just memory size but bandwidth ...

If you take something like a DGX Spark and increase its memory to 512GB that doesn't even solve the problem. Because the bandwidth of DDR5 just can't manage reasonable speeds for decode. If you take a dense model or an MoE model uses up most of that 128GB in active decode you will only get like 15 tok/sec. "Real" datacentre inference boxes use high bandwidth memory that is 10x the speed.

I think we're unfortunately a long way off, unless people learn to accept working with much less intelligent models locally.

The innovation is going to have to come on the research & software side -- we need to find ways to pack more intelligence into a smaller number of parameters.

Re: Kimi-K3 on HuggingFace

#373
from the license:

If the Licensee or any of its affiliates operates a Model as a Service business, and the aggregate revenue of the Licensee and its affiliates exceeds 20 million US dollars (or the equivalent in other currencies) in total over any consecutive 12 months, the Licensee must enter into a separate agreement with Moonshot AI before using the Software or its derivative works for any commercial purpose.

Re: Kimi-K3 on HuggingFace

#374

Earlier quoted context omitted.

"I'd like a car that goes 300mph and gets 100mpg while doing it. I'm aware of a car that gets 100mpg but it is extremely slow." You are describing fundamental tradeoffs. Getting more performance relative to model size and training token amount is what all of the labs are solving.

Labs are focusing on creating models, small or large, that perform well on various benchmarks, including general knowledge, domain-specific expertise, and agentic capabilities. Asking for such a model while wanting to be small and fast would align with what you're describing, which I believe is different from what I'm pointing to. The model I'm describing sacrifices domain knowledge and expertise for agentic reasonin…

[deleted]

Re: Kimi-K3 on HuggingFace

#375

There's no going back on this. This is putting a very capable intelligence in the hands of the masses. Private companies in the US are aching for Trump's protectionism but it'll do nothing. The hardware needed to run this is ofc prohibitive, but actually putting it out there feels like a 'RSA source code on t-shirt' moment for humanity.

a "moment for humanity"? as if this shit isn't going to generate 99% slop at the cost of all we have left as a species?

Sometimes I’m not sure who is more unhinged: the total AI kool aid drinkers who think this will make us all into immortal demigods (or take over the world as it goes “foom”), or the AI doomers and haters who exaggerate everything potentially negative about it and react to it the way a 1980s Christian fundamentalist reacted to rock music.

It’s a new fundamental innovation in math and CS that allows large scale lossy compression of natural language another data formats in a way that is semantically queryable and cross-referenceable. It also manifests some form of emergent intelligence, likely evidence of the long posited link between intelligence and data compression.

The tech is awesome. It’s one of the coolest things I’ve seen in over a decade. The industry is kind of shitty, which is not unusual. The discourse around it is almost universally insane, crazy people arguing with crazy people.

Oh and get off the AI eco bullshit train. Look up the energy cost of AI queries vs driving or running a home air conditioning system. Feel bad about using AI? Skip that DoorDash order. You probably just saved the energy of 1-2 days of heavy Claude Code use.

Re: Kimi-K3 on HuggingFace

#376
From a cursory glance on huggingface, the files don't add up to 2+TB. Unless it adds up to that when you extract the multiple ~17GB files, if that's the case then that's some crazy compression.

Re: Kimi-K3 on HuggingFace

#377
post #360
post #63

Earlier quoted context omitted.

> Even if the output is like 5-6 tok/s, that might be usable for some purposes. You'll spend ~100x more on electricity than the API cost to have it run on someone else's GPU at several hundred tokens per second. I think some sort of extreme data privacy requirement is the only situation that justifies this, but the intersection of {needs absolute data privacy, needs to run SOTA model, cannot afford GPUs} is really re…

One aspect of this is cyberattack proliferation by way of "Hey boss, I saw this TikTok that says if you let me invest [a tiny piece of the neighborhood's profit|our militia's budget] into some RAM, I could get a fully autonomous cyber operation up and running that pays for itself via ransomware etc. within weeks. You like it, we upgrade to something that can work even faster. We don't need the hacker guy from Swordfi…

A similar world is already here.

Young men 14-?? already compromise and attempt to extort organizations daily, sometimes cluelessly from western nations, often not. It doesn’t have to be gangs when the home country doesn’t care / isn’t technologically or culturally developed.

Already seeing AI-written payloads and frameworks in the wild. I think it’ll turn out that AI won’t build you a maintainable ERP but it can create C2 networks, exploit POCs or even 0-days potentially, and let kids make their own ransomware tooling. Then we’re dealing not with a handful of cybercrime tool makers but a generational problem.

Re: Kimi-K3 on HuggingFace

#378

From a cursory glance on huggingface, the files don't add up to 2+TB. Unless it adds up to that when you extract the multiple ~17GB files, if that's the case then that's some crazy compression.

If it's 4-bit native for the sparse parameters (which is the bulk of them) why would you expect it to add up to 2+TB?

Re: Kimi-K3 on HuggingFace

#380

Earlier quoted context omitted.

It will be very interesting to see what kind of 'slow' performance people get from running it on a no GPU, but tons of RAM server (like a dual or quad socket xeon with 1.5 to 3TB of RAM). For the purpose of giving it longer duration tasks to generate a piece of something and come back and check on what it has done in 4 or 6 hours. Even if the output is like 5-6 tok/s, that might be usable for some purposes. Huge pric…

> But I think it's going to need more than 1536GB of RAM, with a usable and large amount of context, more like 2TB and preferably 2.5 to 3TB. The model is known to be MXFP4 according to Kimi's release blog post, so the model weights will be less than 1536GB: https://www.kimi.com/blog/kimi-k3 Also, their previous models were native INT4, so it would be weird if they went larger now.

Update: Looks like the model is larger after all (1561.44 GB). Only the MoE weights are MXFP4, while the other weights are BF16 (and a few FP32).

* Sparse Experts: 1481.4 GB

* Dense Experts: 1.9 GB

* Self-Attention: 72.4 GB

* LLM Head: 2.4 GB

* Embeddings: 2.4 GB

* Vision Encoder: 0.35 GB (surprisingly small)

plus some miscellaneous parameters.

Most importantly, we now know that the model has 104B active parameters, which is quite a lot and will make it difficult to self-host efficiently.

Post reply on HN