Live data from Hacker News

Kimi-K3 on HuggingFace

huggingface.co

251–260 of 588 posts

Re: Kimi-K3 on HuggingFace

#251
post #228

This will be interesting for a few reasons. First, depending on where the median pricing settles w/ 3rd party providers will tell us what it costs to serve a 3T model. Since it's going to be mxfp4 native, it'll take ~1.5TB of VRAM to host this, which is juuust at the limit of 8xb200s (but realistically you'll need 16x for context / throughput optimisation). Won't be cheap to host, but at least we should get some rang…

You're assuming inference providers are going to sell tokens at cost. You're also assuming that the inference providers have will optimized inference engine. I haven't seen that to be the case so far, to be honest. Take a look at GLM 5 vs GLM 5.2 pricing -- GLM 5.2 cost more despite being the same model. Take a look a look at DeepSeek, which hosts DS v4, profitably, yet others aren't able or willing to match the pric…

I think it's unclear the the DS hosted prices are profitable. AFAIK that haven't claimed that.

OTOH, the multiple providers who have settled around the same price point ($3.48/M output tokens for multiple providers with good reputations) does indicate where it is profitable: https://openrouter.ai/deepseek/deepseek-v4-pro#providers

Re: Kimi-K3 on HuggingFace

#252

Earlier quoted context omitted.

Because lack of talent and organizational disfunction matters a lot more than you think. The reason why OAI and Ant are always at the top is because of this and I’d say compute is third on the list.

I would argue that they actually don’t lack talent, they have an insane bench of really smart people. What they lack is any sort of direction and leadership. They are a ship lost in the ocean and up until now have been lucky to find a few treasures along the their way.

> they have an insane bench of really smart people.

Filtered heavily into those who care only about money. Many don’t want to work there. Smart people have other choices.

Re: Kimi-K3 on HuggingFace

#253

Earlier quoted context omitted.

I’ve priced it out: max $135/month to run a dual Xeon 2U server with 3T RAM & 2x 22 core Xeon Gold. It’s the 2x 750W power supplies that ultimately determine opex. My power costs $0.124/kWh, the $135 assumes drawing maximum power continuously, and in that case, I can probably offset my heating bill a little bit in the winter, so maybe effectively a little bit lower. I don’t know if that’s 100x more than I’d pay (opex…

Keep in mind that just because it has dual 750W power supplies that doesn't mean it's what its load will be, for a full CPU loaded wattage figure you'd need basically a pair of kill-a-watts plugged in inline on the feed for each poewr supply and then run stress-ng with artificial cpu stress on all cores for an hour. Under heavy inference load you will find that the cpu usage is actually less as the bottleneck is the…

[deleted]

Re: Kimi-K3 on HuggingFace

#254

Earlier quoted context omitted.

We know labs make money on inference, and we know they lose a lot of money on inference+training.

Just out of curiosity, based on what we know for sure they(OAI+A) make money on pure inference and lose on inference+training?

OpenAI's financials leaked and showed this pretty convincingly.

Anthropic was probably profitable last quarter, without training costs: https://www.wsj.com/tech/ai/mind-blowing-growth-is-about-to-...

Re: Kimi-K3 on HuggingFace

#255
post #48

Earlier quoted context omitted.

> Then we'll be able to guesstimate if "labs are subsidising tokens on API pricing". No, you don't. Without training cost you can infer only the marginal cost of serving this kind of models. Moreover, you don't know the actual size of closed models (what if Fable is a 10T model? What if it's 1T?)

>No, you don't. Without training cost you can infer only the marginal cost of serving this kind of models. Are you talking about Kimi's training cost or the training cost of the model(s) that Kimi distilled? Because Moonshot didn't even incur the majority of the training costs either

> the training cost of the model(s) that Kimi distilled?

The distillation process involves getting conversation traces from the model you are distilling from, and then training your model against them.

You still have to do the training!

Re: Kimi-K3 on HuggingFace

#256

Earlier quoted context omitted.

As a completely one person, single sample anecdote, the 'heretic' uncensored Q8 GGUF variants several people have published of Qwen 3.5-122, 3.6-27B and 3.6-35B-A3B will very happily discuss just about any controversial topic that the CCP hates. Including lots of things that would get you thrown into prison if you published them in Mandarin on the domestic Chinese internet. https://github.com/p-e-w/heretic As a side…

That's a lot of words to say "No, no have has seemingly done that yet with K3".

Yet the comment was valuable nonetheless.

Re: Kimi-K3 on HuggingFace

#258
Are they going to release Kimi K3.1? I’m eager to test it. According to rumors on X, it could outperform Fable. Could it be the first Chinese open-weight model to become the leading frontier model?

Re: Kimi-K3 on HuggingFace

#260
post #136

I heard this is the talk in town these days. Why can't Meta keep up? With >10000000x more resources you'd think that they'd be able to introduce equally performant if not better open weight models

The SemiAnalysis piece on this is long but very much worth reading:

> The company appears burdened by far too many disparate groups that are over-optimizing for certain metrics as opposed to delivering usable technology for the company as a whole.

> And because Meta has a reputation for throwing money at problems and executing at high speed, these U-turns end up becoming more costly versus other companies that take a more disciplined or conservative approach. Suppliers also lose faith when given design wins are later cancelled. This has lead to less supply chain prioritization on new designs. Some suppliers favor focusing on Amazon or Google designs due to Meta’s frequent reshuffling.

> Few inside Meta’s chip division have a full understanding of why the company bought Rivos in the first place, and those who championed the deal internally have since gone quiet.

etc etc

It goes into a lot of depth.

https://newsletter.semianalysis.com/p/metas-infrastructure-t...

Post reply on HN