Live data from Hacker News

Kimi-K3 on HuggingFace

huggingface.co

381–390 of 588 posts

Re: Kimi-K3 on HuggingFace

#382

Earlier quoted context omitted.

The next mac Ultra will allow to run a big model locally with acceptable speed. But we need people to optimize it for that computer, and we’ll be more limited in models we can choose from 128GB is enough to run a large model, quantized, REAPed, with MoE and fast SSD for model weights

Not Kimi K3 large though

They'll just have you buy two studios and thunderbolt them together.

Re: Kimi-K3 on HuggingFace

#384

Earlier quoted context omitted.

On the flip side, I really like being able to inspect its reasoning chain thoroughly, as opposed to the "black box" that Anthropic models are now.

Genuine question, is the reasoning chain different from clicking the status bar under a reply and watching it "think"? Or selecting the "Thinking" transcript view in Claude Code? (both on the desktop app). Seems to me that is very out in the open

That's a summarized and filtered view of the actual reasoning.

OpenAI and Anthropic guard the real reasoning closely. Users have never been able to see it and the API returns an encrypted blob instead of legible reasoning.

Re: Kimi-K3 on HuggingFace

#385
post #228

Earlier quoted context omitted.

You're assuming inference providers are going to sell tokens at cost. You're also assuming that the inference providers have will optimized inference engine. I haven't seen that to be the case so far, to be honest. Take a look at GLM 5 vs GLM 5.2 pricing -- GLM 5.2 cost more despite being the same model. Take a look a look at DeepSeek, which hosts DS v4, profitably, yet others aren't able or willing to match the pric…

I'll be honest, I typed that message while having morning coffee, so it's just a quick reaction from my part, not a heavily researched article in a journal :) But I do think that the median price where this settles will tell us something about the floor at which it is profitable to serve this model. > DeepSeek, which hosts DS v4, profitably I specifically mentioned 3rd party providers, because there can be an argumen…

Based on the best available information, DeepSeek is pricing the API such that they can repay their infra capex over 10 months, while deprecating/amortizing the cost of said infra over 3 years.

For my product, I run GLM 5.2 and other models myself, in production, on rented hardware. Paying API prices would cost much more.

EDIT: You can now see several other third-party providers for Kimi K3 (Nebius, Fireworks). All charge exactly the same as the first-party. Does that mean that their costs are the same? Seems quite unlikely. It's simply not an efficient market, yet.

Re: Kimi-K3 on HuggingFace

#386

Earlier quoted context omitted.

There are a number of use cases where sending the contents of your context and prompts (and the resulting output) to a 3rd party service is off the table as an option, and people will compromise speed for data sovereignty. And not everyone's electricity is equally expensive, I pay about $0.075 USD per kWh. It would for example cost me about $48 a month of electricity (not counting cost of cooling) to run a quad socke…

can send safely context if there’s confidential computing ala my site https://trustedrouter.com/

How do you prove you are running exclusively on Nitro enclave instances or GCP confidential spaces?

Re: Kimi-K3 on HuggingFace

#387
Looks like it's live now, it's approx 17GB per safetensors file x 96 files, so so about 1.63TB. I can only imagine that people with their favorite quantizing tools warmed up and ready to go are aggressively downloading it now.

Re: Kimi-K3 on HuggingFace

#388
post #385

Earlier quoted context omitted.

I'll be honest, I typed that message while having morning coffee, so it's just a quick reaction from my part, not a heavily researched article in a journal :) But I do think that the median price where this settles will tell us something about the floor at which it is profitable to serve this model. > DeepSeek, which hosts DS v4, profitably I specifically mentioned 3rd party providers, because there can be an argumen…

Based on the best available information, DeepSeek is pricing the API such that they can repay their infra capex over 10 months, while deprecating/amortizing the cost of said infra over 3 years. For my product, I run GLM 5.2 and other models myself, in production, on rented hardware. Paying API prices would cost much more. EDIT: You can now see several other third-party providers for Kimi K3 (Nebius, Fireworks). All c…

They need to get a license from moonshot to provide inference for K3. Probably have to follow the pricing set by moonshot as well.

Re: Kimi-K3 on HuggingFace

#389
post #75

Earlier quoted context omitted.

There are a number of use cases where sending the contents of your context and prompts (and the resulting output) to a 3rd party service is off the table as an option, and people will compromise speed for data sovereignty. And not everyone's electricity is equally expensive, I pay about $0.075 USD per kWh. It would for example cost me about $48 a month of electricity (not counting cost of cooling) to run a quad socke…

Great, so the other member of the set matters for you more than cost. Do you actually need to run the state of art model at 5 tokens per second instead of a qwen or whatever 7b or 30b model at 100 tokens per second?

The whole mentality of thinking one knows better than another about what they need causes infinitely more problems than it solves.

Re: Kimi-K3 on HuggingFace

#390

from the license: If the Licensee or any of its affiliates operates a Model as a Service business, and the aggregate revenue of the Licensee and its affiliates exceeds 20 million US dollars (or the equivalent in other currencies) in total over any consecutive 12 months, the Licensee must enter into a separate agreement with Moonshot AI before using the Software or its derivative works for any commercial purpose.

good find! This sounds a bit like what Meta was doing with the earlier Llama models?

There is also this paragraph in their licence that is smart marketing-wise:

> 3. If the Software (or any derivative works thereof) is used for any of the Licensee's commercial products or services that have more than 100 million monthly active users, or more than 20 million US dollars (or equivalent in other currencies) in monthly revenue, "Kimi K3" must be prominently displayed on the user interface of such product or service.

Post reply on HN