Many people are talking about price, but I think that the most interesting aspect of this release, by far, is customization. Any startup can download the weights, tinker with them, and fine-tune them. The real win here isn't necessarily cost, but performance on your data and IP sovereignty. It's a huge win. Kudos to the Kimi team.
Kimi-K3 on HuggingFace
561–570 of 588 posts
Re: Kimi-K3 on HuggingFace
#562from the license: If the Licensee or any of its affiliates operates a Model as a Service business, and the aggregate revenue of the Licensee and its affiliates exceeds 20 million US dollars (or the equivalent in other currencies) in total over any consecutive 12 months, the Licensee must enter into a separate agreement with Moonshot AI before using the Software or its derivative works for any commercial purpose.
But then how thick does that layer need to be?
I think the real win from these models is enabling well-funded/profitable companies to not be beholden to the big closed-weight providers. But I like the idea of letting a hundred flowers bloom for under $20M/y each.
Re: Kimi-K3 on HuggingFace
#563It is online on https://app.fireworks.ai/models/fireworks/kimi-k3 (Uncached Input $3.00/M Cached Input $0.30/M Output $15.00/M)
Fireworks' priority tier of Kimi (at $3.75/M vs. Moonshot's $3.00/M) is available on OpenRouter as well. https://openrouter.ai/moonshotai/kimi-k3#providers Currently it's showing significantly better latency, but at a fraction of the usage Moonshot is experiencing, so we'll see how that holds up - regardless, a same-day deployment is an impressive feat!
Re: Kimi-K3 on HuggingFace
#564I feel like most hardware to run LLMs on is shaped wrong for individuals. It's either having a model struggling along with like 5-10 tokens per second on unified memory, or data center cards with hundreds of GB of VRAM consuming more than a kW of power. It doesn't seem like there's prosumer GPUs with like 180W-250W TDP and 128 GB or 256 GB of VRAM (one can dream). Then bifurcation and even just two of those cards wou…
https://www.youtube.com/watch?v=nbHOBvLlypY https://www.youtube.com/watch?v=PV89U-PNUNA
Re: Kimi-K3 on HuggingFace
#565Re: Kimi-K3 on HuggingFace
#566Many people are talking about price, but I think that the most interesting aspect of this release, by far, is customization. Any startup can download the weights, tinker with them, and fine-tune them. The real win here isn't necessarily cost, but performance on your data and IP sovereignty. It's a huge win. Kudos to the Kimi team.
Who does this, though? I'm truly curious.
Re: Kimi-K3 on HuggingFace
#567Earlier quoted context omitted.
That's an unusually low electric rate for the US - way below the lowest state average which is Idaho at 12.4 cents. It's certainly possible that you are getting 7.5 cents including delivery, but I've had friends say that they're "getting 13 cents per kWh" here in Massachusetts, but that's just the supply rate and the delivery is another ~18 cents. There are parts of states like Grant County Washington that have cheap…
I pay a little less than that though I’m limited to 50 kW average load. Delivery is a fixed monthly fee, about $10 USD/month.
Re: Kimi-K3 on HuggingFace
#568Re: Kimi-K3 on HuggingFace
#569I heard this is the talk in town these days. Why can't Meta keep up? With >10000000x more resources you'd think that they'd be able to introduce equally performant if not better open weight models
Re: Kimi-K3 on HuggingFace
#570Earlier quoted context omitted.
It will be very interesting to see what kind of 'slow' performance people get from running it on a no GPU, but tons of RAM server (like a dual or quad socket xeon with 1.5 to 3TB of RAM). For the purpose of giving it longer duration tasks to generate a piece of something and come back and check on what it has done in 4 or 6 hours. Even if the output is like 5-6 tok/s, that might be usable for some purposes. Huge pric…
> Even if the output is like 5-6 tok/s, that might be usable for some purposes. You'll spend ~100x more on electricity than the API cost to have it run on someone else's GPU at several hundred tokens per second. I think some sort of extreme data privacy requirement is the only situation that justifies this, but the intersection of {needs absolute data privacy, needs to run SOTA model, cannot afford GPUs} is really re…
this is hacker news sir
we don't do things because they are easy ... or because they are justified ...