Live data from Hacker News

Kimi-K3 on HuggingFace

huggingface.co

221–230 of 588 posts

Re: Kimi-K3 on HuggingFace

#221

Earlier quoted context omitted.

LLM inference unfortunately also seems to be a task that's poorly formed for moderate consumer hardware,as a single user. For a single user use case, the load is bursty but requires the weights to be in memory already. So a multi user server that keeps the model weights in parts of its memory and then spends some more per user kv cache is wildly more efficient and the wildly expensive gpu cores aren't just sitting id…

Why does it have to be so bursty though? Just let it run multiple continuous-batched inferences overnight. This would work especially well in combination with SSD offload, and given any kind of sparse attention (common in more recent models) even swapping out the KV cache itself to disk might ultimately be a win. I wouldn't be surprised if something like that ultimately became feasible for single users running even K…

If your workload fits long batches throughout the night you could schedule them better, yeah. But I think very few have a usage pattern like that?

Given the hardware shortage in the world, I suspect renting ("sharing") via APIs will likely remain cheaper for the foreseeable future since each piece of hardware isn't sitting idle nearly as much.

Re: Kimi-K3 on HuggingFace

#223

Earlier quoted context omitted.

Do you mean by trading dollars for the privacy you need as: a) Contracting with a third-party independent inference provider who will run your choice of model on fast hardware that they own, with all appropriate data security/privacy/contractual/compliance protection in place or b) Contracting with the original creators of the model to run inference via their API and with assurances that all the same data protection…

A) is very doable with e.g. Amazon Bedrock. They'll give you HIPAA compliance, they even have a data center for US government classified data, they can give you European data sovereignty. And with OpenAI and Anthropic models to boot, you don't even have to settle for open weights. What kind of privacy needs do you really have beyond that?

It is not my use case but given recent political developments in international relations caused by the executive branch of the US government, off the top of my head, I could think of a lot of European or Canadian firms for which that would not be an option. No matter what they might promise about European sovereignty. For a good 'ol patriotic US domestic company? Sure.

Re: Kimi-K3 on HuggingFace

#224
post #207
post #203

Earlier quoted context omitted.

You can make the same argument for closed models. Why spend hundreds of billions training larger and larger models when you can just use Chinese models? Spend that money somewhere else further up the stack where there’s more value. Let China do the training since they’re so efficient at it.

as a big tech company you have the resources to make many bets and do a lot of things at the same time. it's good to have some specialists with knowledge of model training "just in case".

Even Meta, with all their resources, makes all their hardware in China. Manufacturing anywhere else is just burning cash.

Re: Kimi-K3 on HuggingFace

#225
post #31

We already know that competition brought GLM 5.2 prices down roughly 45% since its release on June 16th (1.5 months ago), and the price downward slope is probably still going (I've been checking regularly and new providers keep fighting on price, I don't think prices have settled yet). For reference : https://openrouter.ai/z-ai/glm-5.2#providers I saw arguments like "Providers cannot price less than their costs" in o…

> My guess is that someone will be selling tokens at less than electricity + depreciation of GCs soon, since there's a lot of competition and "smaller" data centers have overcapacity? This is speculation, correct me if I'm wrong My guess is they are selling you the tokens, then selling your tokens (data) onto someone else.

I see these conspiratorial arguments all the time and I think people massively overestimate the value of the average users tokens.

The problems with frontier models (design taste, ability to solve novel/difficult problems, etc) cannot be solved by throwing more slop from the average user at it.

Actually, most of the main deficiencies in current models stem from the fact that their data sets aren’t curated and specialized enough.

Re: Kimi-K3 on HuggingFace

#226

Earlier quoted context omitted.

I dunno, K3 thinks a lot before it actually replies, and you might be in the ~1 tok/speed region or even "seconds / tokens", and with K3, you'd wait days if not weeks for a reply in that case. Don't get me wrong, slow is sometimes better than "not at all", but depending on the performance, it might end up way too slow to even work for batched/async jobs like that.

I agree it's very likely to be painfully slow, I very much want to see some real world results from people who try it. Early testers will inform others on whether it's even worth trying. Results very much TBD right now. I don't have a system sitting around here with 2TB of greater of RAM that isn't already committed for other uses, regretfully.

You can rent one in the cloud to try it

Re: Kimi-K3 on HuggingFace

#227
post #214

Earlier quoted context omitted.

I agree but worth noting that it's never gonna be very practical to run LLMs like this at home. Unless we have some sort of design breakthrough, the only "sensible" way to run them is at high batch levels on shared HW. Like, yeah if I could spend a few grand on such a GPU I probably would coz I'm a rich nerd, but I'd acknowledge it as an extremely inefficient luxury, kinda like a sports car. So I think you could say…

"Never" is a long time. Just think about how much ram we had 10 or 20 years ago. 1.5TB isn't a lot really.

The typical ram has surprisingly not increased very much in 10 years.

> April 2016, 8 GB was standard across the 13-inch MacBook Air range

... Now it's 16.

Rich nerds will have quite a bit more. But I suspect the standard of model rich nerds want to use will have gone up somewhat too.

Re: Kimi-K3 on HuggingFace

#228

This will be interesting for a few reasons. First, depending on where the median pricing settles w/ 3rd party providers will tell us what it costs to serve a 3T model. Since it's going to be mxfp4 native, it'll take ~1.5TB of VRAM to host this, which is juuust at the limit of 8xb200s (but realistically you'll need 16x for context / throughput optimisation). Won't be cheap to host, but at least we should get some rang…

You're assuming inference providers are going to sell tokens at cost. You're also assuming that the inference providers have will optimized inference engine. I haven't seen that to be the case so far, to be honest.

Take a look at GLM 5 vs GLM 5.2 pricing -- GLM 5.2 cost more despite being the same model.

Take a look a look at DeepSeek, which hosts DS v4, profitably, yet others aren't able or willing to match the price.

Re: Kimi-K3 on HuggingFace

#229
post #48

Earlier quoted context omitted.

> Then we'll be able to guesstimate if "labs are subsidising tokens on API pricing". No, you don't. Without training cost you can infer only the marginal cost of serving this kind of models. Moreover, you don't know the actual size of closed models (what if Fable is a 10T model? What if it's 1T?)

> No, you don't. Without training cost you can infer only the marginal cost of serving this kind of models. Still useful; "are the labs marginally profitable just on the marginal inference costs?" is still a useful question to answer. After all, if they aren't even profitable on inference in isolation, then we can expect to see large price increases. If they are able to turn a marginal profit on inference alone, then…

We know labs make money on inference, and we know they lose a lot of money on inference+training.

Re: Kimi-K3 on HuggingFace

#230

There's no going back on this. This is putting a very capable intelligence in the hands of the masses. Private companies in the US are aching for Trump's protectionism but it'll do nothing. The hardware needed to run this is ofc prohibitive, but actually putting it out there feels like a 'RSA source code on t-shirt' moment for humanity.

No, luckily private companies in the US are aching for this to NOT be banned. NVIDIA, Microsoft, etc. just released that letter. We’re saved from the trillionaire companies (OpenAI, Anthropic) by the other trillionaire companies acting in self-interest (hosting and hardware).
Post reply on HN