Live data from Hacker News

Kimi-K3 on HuggingFace

huggingface.co

561–570 of 588 posts

Re: Kimi-K3 on HuggingFace

#561

Many people are talking about price, but I think that the most interesting aspect of this release, by far, is customization. Any startup can download the weights, tinker with them, and fine-tune them. The real win here isn't necessarily cost, but performance on your data and IP sovereignty. It's a huge win. Kudos to the Kimi team.

Who does this, though? I'm truly curious.

Re: Kimi-K3 on HuggingFace

#562

from the license: If the Licensee or any of its affiliates operates a Model as a Service business, and the aggregate revenue of the Licensee and its affiliates exceeds 20 million US dollars (or the equivalent in other currencies) in total over any consecutive 12 months, the Licensee must enter into a separate agreement with Moonshot AI before using the Software or its derivative works for any commercial purpose.

If you put a layer between the customer and the LLM, a la ChatGPT/Claude, are you a Model-as-a-Service business? I think you are not.

But then how thick does that layer need to be?

I think the real win from these models is enabling well-funded/profitable companies to not be beholden to the big closed-weight providers. But I like the idea of letting a hundred flowers bloom for under $20M/y each.

Re: Kimi-K3 on HuggingFace

#563
post #407

It is online on https://app.fireworks.ai/models/fireworks/kimi-k3 (Uncached Input $3.00/M Cached Input $0.30/M Output $15.00/M)

Fireworks' priority tier of Kimi (at $3.75/M vs. Moonshot's $3.00/M) is available on OpenRouter as well. https://openrouter.ai/moonshotai/kimi-k3#providers Currently it's showing significantly better latency, but at a fraction of the usage Moonshot is experiencing, so we'll see how that holds up - regardless, a same-day deployment is an impressive feat!

Self-replying as this is no longer correct - there are two tiers, one matching Moonshot pricing with comparable latency/throughput, and a fast mode at $4.50/M (still cheaper than Opus) with 3x the throughput.

Re: Kimi-K3 on HuggingFace

#564

I feel like most hardware to run LLMs on is shaped wrong for individuals. It's either having a model struggling along with like 5-10 tokens per second on unified memory, or data center cards with hundreds of GB of VRAM consuming more than a kW of power. It doesn't seem like there's prosumer GPUs with like 180W-250W TDP and 128 GB or 256 GB of VRAM (one can dream). Then bifurcation and even just two of those cards wou…

You can run GLM 5.2 at a reasonable quant on a cluster of DGX Sparks. On the link below, the guy is getting decent performance and links to a repo. An 8x Spark cluster can run it even better. A Spark cluster is about as good as it gets for (a) runnable at home, (b) "affordable" hardware, (c) not going to kill your power bill.

https://www.youtube.com/watch?v=nbHOBvLlypY https://www.youtube.com/watch?v=PV89U-PNUNA

Re: Kimi-K3 on HuggingFace

#566

Many people are talking about price, but I think that the most interesting aspect of this release, by far, is customization. Any startup can download the weights, tinker with them, and fine-tune them. The real win here isn't necessarily cost, but performance on your data and IP sovereignty. It's a huge win. Kudos to the Kimi team.

Who does this, though? I'm truly curious.

[deleted]

Re: Kimi-K3 on HuggingFace

#567
post #309

Earlier quoted context omitted.

That's an unusually low electric rate for the US - way below the lowest state average which is Idaho at 12.4 cents. It's certainly possible that you are getting 7.5 cents including delivery, but I've had friends say that they're "getting 13 cents per kWh" here in Massachusetts, but that's just the supply rate and the delivery is another ~18 cents. There are parts of states like Grant County Washington that have cheap…

I pay a little less than that though I’m limited to 50 kW average load. Delivery is a fixed monthly fee, about $10 USD/month.

To me (EU) that seems like a pretty generous limit. A normal house connection here is limited to 8kW single phase or 17kW for a 3 phase connection. You can get more, but that is very uncommon and gets expensive quickly

Re: Kimi-K3 on HuggingFace

#569
post #136

I heard this is the talk in town these days. Why can't Meta keep up? With >10000000x more resources you'd think that they'd be able to introduce equally performant if not better open weight models

I heard Muse Spark 1.1 isn't too shabby. Since Meta compute >> Chinese labs, we'll see how it goes from here.

Re: Kimi-K3 on HuggingFace

#570
post #63

Earlier quoted context omitted.

It will be very interesting to see what kind of 'slow' performance people get from running it on a no GPU, but tons of RAM server (like a dual or quad socket xeon with 1.5 to 3TB of RAM). For the purpose of giving it longer duration tasks to generate a piece of something and come back and check on what it has done in 4 or 6 hours. Even if the output is like 5-6 tok/s, that might be usable for some purposes. Huge pric…

> Even if the output is like 5-6 tok/s, that might be usable for some purposes. You'll spend ~100x more on electricity than the API cost to have it run on someone else's GPU at several hundred tokens per second. I think some sort of extreme data privacy requirement is the only situation that justifies this, but the intersection of {needs absolute data privacy, needs to run SOTA model, cannot afford GPUs} is really re…

> that justifies this

this is hacker news sir

we don't do things because they are easy ... or because they are justified ...

Post reply on HN