Live data from Hacker News

Kimi-K3 on HuggingFace

huggingface.co

481–490 of 588 posts

Re: Kimi-K3 on HuggingFace

#481

I suggest downloading these frontier models just to have a copy; even though it’s 1.5TB, it’s worth sticking in a cheap disk and putting aside. Seeding torrents would be even more useful. The man is coming to lock these down, like they tried to do with encryption algorithms. The only way open software survives regulation is through distribution. Over time the enormous investment in techniques and hardware manufacturi…

they will just restrict you from buying the hardware these run on

The rising price of hardware is already doing just that

Re: Kimi-K3 on HuggingFace

#482
post #12

Did someone run censorship and political bias tests on this ? Must be interesting.

Outside of asking it to talk about Tiananmen Square, are there any standard tests for "bias"? And if so, who created them and what are their biases?

Kimi is basically a Chinese nationalist in every sense. Do I trust it to not insert software backdoors?

Re: Kimi-K3 on HuggingFace

#483

It is online on https://app.fireworks.ai/models/fireworks/kimi-k3 (Uncached Input $3.00/M Cached Input $0.30/M Output $15.00/M)

Yep, it’s also availability on Together.ai for your sale price. Fireworks and Together were the first places I checked!

Re: Kimi-K3 on HuggingFace

#484
post #291

Earlier quoted context omitted.

Say a single Kimi K3 is deployed on 16 x B200s: how many concurrent users can that handle? I realize the question assumes a major simplification that everyone's prompts/sessions are the same.

>I realize the question assumes a major simplification that everyone's prompts/sessions are the same. well, exactly. that's tough to answer without just average sampling because some users will ask the model "what's todays date" or "what color is the sky?" and some users will ask "Let's rewrite the linux kernel in brainfuck."

So maybe the better question would be how many concurrent actively-thinking/working agents it could handle?

Re: Kimi-K3 on HuggingFace

#485

Earlier quoted context omitted.

good find! This sounds a bit like what Meta was doing with the earlier Llama models? There is also this paragraph in their licence that is smart marketing-wise: > 3. If the Software (or any derivative works thereof) is used for any of the Licensee's commercial products or services that have more than 100 million monthly active users, or more than 20 million US dollars (or equivalent in other currencies) in monthly re…

Is that even enforceable?

Before figuring that out, could Facebook take you to court in order to argue their case that it is enforceable, and thereby forcing you to get lawyers and be distracted by the preparation and all that comes with this?

Re: Kimi-K3 on HuggingFace

#486

I feel like most hardware to run LLMs on is shaped wrong for individuals. It's either having a model struggling along with like 5-10 tokens per second on unified memory, or data center cards with hundreds of GB of VRAM consuming more than a kW of power. It doesn't seem like there's prosumer GPUs with like 180W-250W TDP and 128 GB or 256 GB of VRAM (one can dream). Then bifurcation and even just two of those cards wou…

LLM inference unfortunately also seems to be a task that's poorly formed for moderate consumer hardware,as a single user. For a single user use case, the load is bursty but requires the weights to be in memory already. So a multi user server that keeps the model weights in parts of its memory and then spends some more per user kv cache is wildly more efficient and the wildly expensive gpu cores aren't just sitting id…

A decentralized inference network would be cool. Something that's set up so that I can run a model for personal use on beefy hardware, but also farm out the unused GPU time to the network, probably at much lower prices than normal providers since it would be slower and would lack data security guarantees.

Re: Kimi-K3 on HuggingFace

#487
post #309

Earlier quoted context omitted.

That's an unusually low electric rate for the US - way below the lowest state average which is Idaho at 12.4 cents. It's certainly possible that you are getting 7.5 cents including delivery, but I've had friends say that they're "getting 13 cents per kWh" here in Massachusetts, but that's just the supply rate and the delivery is another ~18 cents. There are parts of states like Grant County Washington that have cheap…

A lot of people quoting low rates are also just referring to their off-peak rate. This is pretty common in EV discussions. It's not exactly a fair argument there, either, because the flip side of having an off-peak rate is that the on-peak rate is usually quite a lot higher. So the true effective rate is a bit higher, somewhere in the middle depending on usage pattern.

[deleted]

Re: Kimi-K3 on HuggingFace

#488
post #63

Earlier quoted context omitted.

> Even if the output is like 5-6 tok/s, that might be usable for some purposes. You'll spend ~100x more on electricity than the API cost to have it run on someone else's GPU at several hundred tokens per second. I think some sort of extreme data privacy requirement is the only situation that justifies this, but the intersection of {needs absolute data privacy, needs to run SOTA model, cannot afford GPUs} is really re…

There are a number of use cases where sending the contents of your context and prompts (and the resulting output) to a 3rd party service is off the table as an option, and people will compromise speed for data sovereignty. And not everyone's electricity is equally expensive, I pay about $0.075 USD per kWh. It would for example cost me about $48 a month of electricity (not counting cost of cooling) to run a quad socke…

>and people will compromise speed for data sovereignty

People should always compromise speed for data sovereignty! Who said: that in this digital day and age, information about money is more important than money!

Re: Kimi-K3 on HuggingFace

#489

Earlier quoted context omitted.

Do you mean by trading dollars for the privacy you need as: a) Contracting with a third-party independent inference provider who will run your choice of model on fast hardware that they own, with all appropriate data security/privacy/contractual/compliance protection in place or b) Contracting with the original creators of the model to run inference via their API and with assurances that all the same data protection…

A) is very doable with e.g. Amazon Bedrock. They'll give you HIPAA compliance, they even have a data center for US government classified data, they can give you European data sovereignty. And with OpenAI and Anthropic models to boot, you don't even have to settle for open weights. What kind of privacy needs do you really have beyond that?

It's also worth considering what you are actually paying for. And it's not keeping the data private, it's taking the blame when there is a breach. Same reason companies hire big consulting firms whenever they need to make an important but possibly risky decision.

Re: Kimi-K3 on HuggingFace

#490

I suggest downloading these frontier models just to have a copy; even though it’s 1.5TB, it’s worth sticking in a cheap disk and putting aside. Seeding torrents would be even more useful. The man is coming to lock these down, like they tried to do with encryption algorithms. The only way open software survives regulation is through distribution. Over time the enormous investment in techniques and hardware manufacturi…

they will just restrict you from buying the hardware these run on

[deleted]
Post reply on HN