Live data from Hacker News

Kimi-K3 on HuggingFace

huggingface.co

411–420 of 588 posts

Re: Kimi-K3 on HuggingFace

#411
post #407

It is online on https://app.fireworks.ai/models/fireworks/kimi-k3 (Uncached Input $3.00/M Cached Input $0.30/M Output $15.00/M)

Fireworks' priority tier of Kimi (at $3.75/M vs. Moonshot's $3.00/M) is available on OpenRouter as well. https://openrouter.ai/moonshotai/kimi-k3#providers Currently it's showing significantly better latency, but at a fraction of the usage Moonshot is experiencing, so we'll see how that holds up - regardless, a same-day deployment is an impressive feat!

I've used GLM-5.2 a lot on fireworks and had never ever issues on rate limits. If they cannot handle the load with K3, there's the priority tier to get your evals done.

I'm definitely having full eval suite on already if they get overloaded later on.

Re: Kimi-K3 on HuggingFace

#412

Earlier quoted context omitted.

can send safely context if there’s confidential computing ala my site https://trustedrouter.com/

How do you prove you are running exclusively on Nitro enclave instances or GCP confidential spaces?

Is your local compute airgapped?

Re: Kimi-K3 on HuggingFace

#414

Earlier quoted context omitted.

Having worked in / adjacent several such industries, a lot of the question depends on scale. A trillion-dollar business can easily trade dollars for the privacy. A business with $1M to spend won't even get a phone call with OpenAI or Anthropic, who were the only* previous players in town for doing this. Worst-case example: Bootstrapped startup working in military. It's also the case that an open model enables many mo…

> Worst-case example: Bootstrapped startup working in military. That's the easiest case. AWS Bedrock models running in AWS Secret Cloud for Industry. (I really have no affiliation with them, I'm just like... this is a completely solved problem, why do people think this is hard and requires on-prem hardware?) https://www.aboutamazon.com/news/aws/aws-secret-cloud-for-in... I'm with GP that these are tinfoil hat concern…

You seem to be categorizing everything that considers their data being in the possession of the US an unacceptable risk to be tinfoil hat, which is kind of an insult to a large portion of the world. If you haven't been paying attention to the news in the last 48 months, the political reality has shifted considerably.

Note that the other commenter never said US-based military oriented startup. You just assumed, then jumped to "heck yeah let's use Amazon Secret Cloud for Industry"

Not everyone has or wants an office in Crystal City.

Re: Kimi-K3 on HuggingFace

#415
post #309

Earlier quoted context omitted.

There are a number of use cases where sending the contents of your context and prompts (and the resulting output) to a 3rd party service is off the table as an option, and people will compromise speed for data sovereignty. And not everyone's electricity is equally expensive, I pay about $0.075 USD per kWh. It would for example cost me about $48 a month of electricity (not counting cost of cooling) to run a quad socke…

That's an unusually low electric rate for the US - way below the lowest state average which is Idaho at 12.4 cents. It's certainly possible that you are getting 7.5 cents including delivery, but I've had friends say that they're "getting 13 cents per kWh" here in Massachusetts, but that's just the supply rate and the delivery is another ~18 cents. There are parts of states like Grant County Washington that have cheap…

A lot of people quoting low rates are also just referring to their off-peak rate. This is pretty common in EV discussions. It's not exactly a fair argument there, either, because the flip side of having an off-peak rate is that the on-peak rate is usually quite a lot higher. So the true effective rate is a bit higher, somewhere in the middle depending on usage pattern.

Re: Kimi-K3 on HuggingFace

#416

Earlier quoted context omitted.

That's very interesting. Does that mean you can reduce say, a 30B class Q8 from ~30 GB down to 10 GB or less?

704gb -> 564gb; 358 gb -> 270 gb; 28.79 gb -> 7.65 gb; 439 gb -> 93 gb It depends on the total entropy of the model. Smaller models have less entropy.

> Smaller models have less entropy.

Interesting. Why is that? I would have expected the opposite, since larger models have to try less hard to fit the training data. Or maybe this leaves more parameters with random initialization, resulting in higher entropy for larger models?

Re: Kimi-K3 on HuggingFace

#417

Earlier quoted context omitted.

There are a number of use cases where sending the contents of your context and prompts (and the resulting output) to a 3rd party service is off the table as an option, and people will compromise speed for data sovereignty. And not everyone's electricity is equally expensive, I pay about $0.075 USD per kWh. It would for example cost me about $48 a month of electricity (not counting cost of cooling) to run a quad socke…

There are a number of use cases where sending the contents of your context and prompts (and the resulting output) to a 3rd party service is off the table as an option, and people will compromise speed for data sovereignty. Are there? At the highest levels of defense and law, AWS and Azure are used. Having tried selling some of these entities on doing things in-house, there seems to be little interest.

> Are there? At the highest levels of defense and law, AWS and Azure are used.

This is certainly true if the user is an American company. You could look at the European initiatives to run this stuff on hardware they own in facilities they own and control within the borders of Europe for a counter-example.

Such as: https://www.google.com/search?client=firefox-b-d&q=schwarz+s...

https://www.dutchnews.nl/2026/04/government-turns-to-german-...

Re: Kimi-K3 on HuggingFace

#418

This will be interesting for a few reasons. First, depending on where the median pricing settles w/ 3rd party providers will tell us what it costs to serve a 3T model. Since it's going to be mxfp4 native, it'll take ~1.5TB of VRAM to host this, which is juuust at the limit of 8xb200s (but realistically you'll need 16x for context / throughput optimisation). Won't be cheap to host, but at least we should get some rang…

>realistically you'll need 16x for context / throughput optimisation

Sounds like I'm buying a lottery ticket this week so I can drop $800k on hardware.

Re: Kimi-K3 on HuggingFace

#419
post #384

Earlier quoted context omitted.

Genuine question, is the reasoning chain different from clicking the status bar under a reply and watching it "think"? Or selecting the "Thinking" transcript view in Claude Code? (both on the desktop app). Seems to me that is very out in the open

That's a summarized and filtered view of the actual reasoning. OpenAI and Anthropic guard the real reasoning closely. Users have never been able to see it and the API returns an encrypted blob instead of legible reasoning.

Older models did show the full unredacted thinking trace, but I don't think Opus has ever shown full CoT.

Here is an archived version of Anthropic's API docs saying that Sonnet 3.7 (only) has unredacted CoT on API: https://web.archive.org/web/20260324051339/https://platform....

Re: Kimi-K3 on HuggingFace

#420

Earlier quoted context omitted.

They are saying that AMD's new Epyc Venice CPU has 16 memory channels allowing up to 1.6Tb/s of bandwidth. Which is higher bandwidth than most non-HBM GPUs. So full CPU local AI inference may become viable option in coming years.

This is essentially guaranteed. There are lots of useful smaller models that we should be able to run locally. Over time they'll be more and more capable and require less API usage.

Im wondering if we are finally seeing the end of the "hard disk" era, and are entering a new era of vast instant on systems.
Post reply on HN