Live data from Hacker News

Kimi-K3 on HuggingFace

huggingface.co

431–440 of 588 posts

Re: Kimi-K3 on HuggingFace

#431

Earlier quoted context omitted.

A lot of people quoting low rates are also just referring to their off-peak rate. This is pretty common in EV discussions. It's not exactly a fair argument there, either, because the flip side of having an off-peak rate is that the on-peak rate is usually quite a lot higher. So the true effective rate is a bit higher, somewhere in the middle depending on usage pattern.

It's easy to have your EV only charge off-peak, though. It's just a setting.

My point is that the tradeoff to get off-peak pricing is that on-peak is way, way more expensive. So you can charge the EV off-peak to maximize the savings, but everything else you do during on-peak time costs way more.

Using myself as an example:

I adjust my A/C to run outside of 5pm-9pm (peak) if at all possible, we try to avoid pointless high-draw usage during that same window, and both of our EVs hold off charging until after 9pm.

My rate from 5pm-9pm is 0.43/kWh. My rate after 9pm is 0.09/kWh. The flat rate alternative, if I did not want to worry about time of day, would be 0.21/kWh. These prices are all-in, including transmission and distribution/whatever.

It would be dishonest to say that my EVs only cost me 0.09/kWh to operate, which on it's face is a claim to paying over 50% less. In reality, time of day pricing typically saves me somewhere between 10% and 15% in an average month compared with flat rate.

Re: Kimi-K3 on HuggingFace

#432
post #116

Earlier quoted context omitted.

There's an emerging practice of using Q4 quants and Q8 KV cache for local inference. At that point you can run both Qwen3.5-122B-A10B (my personal choice on Framework Desktop 128gb) and Laguna-S-2.1. Now whether that's good enough for one's use-case remains to be determined. You can also get more out of those (local models and quantizations) if you further tweak the harness you use them with, but tbh this is where it…

Artificial Analysis ranks qwen3.6-27b higher than qwen3.5-122b-a10b on both intelligence and coding. Does that run counter to your experience?

This tracks, in my experience the 27B is better at coding and instruction following. I'm shocked at how much of a difference the dense models vs MoE makes.

But it's a moot point, because for local inference on consumer hardware, the MoE is so much faster.

Re: Kimi-K3 on HuggingFace

#433

Earlier quoted context omitted.

How do you prove you are running exclusively on Nitro enclave instances or GCP confidential spaces?

Is your local compute airgapped?

I trust my local compute quite a bit more than a random project that has "trusted" in its name.

Re: Kimi-K3 on HuggingFace

#434
post #309

Earlier quoted context omitted.

There are a number of use cases where sending the contents of your context and prompts (and the resulting output) to a 3rd party service is off the table as an option, and people will compromise speed for data sovereignty. And not everyone's electricity is equally expensive, I pay about $0.075 USD per kWh. It would for example cost me about $48 a month of electricity (not counting cost of cooling) to run a quad socke…

That's an unusually low electric rate for the US - way below the lowest state average which is Idaho at 12.4 cents. It's certainly possible that you are getting 7.5 cents including delivery, but I've had friends say that they're "getting 13 cents per kWh" here in Massachusetts, but that's just the supply rate and the delivery is another ~18 cents. There are parts of states like Grant County Washington that have cheap…

If you run off solar with battery backup, you can achieve lower than those rates! Look at Time of Use rates. The super off peak rates instantly become the max price point once you pair TOU with Solar + battery.

Re: Kimi-K3 on HuggingFace

#435

Earlier quoted context omitted.

It's easy to have your EV only charge off-peak, though. It's just a setting.

My point is that the tradeoff to get off-peak pricing is that on-peak is way, way more expensive. So you can charge the EV off-peak to maximize the savings, but everything else you do during on-peak time costs way more. Using myself as an example: I adjust my A/C to run outside of 5pm-9pm (peak) if at all possible, we try to avoid pointless high-draw usage during that same window, and both of our EVs hold off chargin…

[deleted]

Re: Kimi-K3 on HuggingFace

#436

Earlier quoted context omitted.

good find! This sounds a bit like what Meta was doing with the earlier Llama models? There is also this paragraph in their licence that is smart marketing-wise: > 3. If the Software (or any derivative works thereof) is used for any of the Licensee's commercial products or services that have more than 100 million monthly active users, or more than 20 million US dollars (or equivalent in other currencies) in monthly re…

Is that even enforceable?

I think it is as enforceable as other licenses are.

Re: Kimi-K3 on HuggingFace

#437
So, in normal parlance this is a 2.8T-A104B model at MXFP4 (weight) * MXFP8 (activations)

Perhaps let's call it Kimi-K3-2.8T-A104B to make matters clear.

Re: Kimi-K3 on HuggingFace

#438

Earlier quoted context omitted.

Do you mean by trading dollars for the privacy you need as: a) Contracting with a third-party independent inference provider who will run your choice of model on fast hardware that they own, with all appropriate data security/privacy/contractual/compliance protection in place or b) Contracting with the original creators of the model to run inference via their API and with assurances that all the same data protection…

A) is very doable with e.g. Amazon Bedrock. They'll give you HIPAA compliance, they even have a data center for US government classified data, they can give you European data sovereignty. And with OpenAI and Anthropic models to boot, you don't even have to settle for open weights. What kind of privacy needs do you really have beyond that?

>very doable with

for anyone not US-based, this company is hostile and you have to assume the US government can and will force them to give access to your data.

Re: Kimi-K3 on HuggingFace

#439

This will be interesting for a few reasons. First, depending on where the median pricing settles w/ 3rd party providers will tell us what it costs to serve a 3T model. Since it's going to be mxfp4 native, it'll take ~1.5TB of VRAM to host this, which is juuust at the limit of 8xb200s (but realistically you'll need 16x for context / throughput optimisation). Won't be cheap to host, but at least we should get some rang…

AISI is capped at 100M tokens and K3 is less token efficient than Anthropic/OpenAI models. There is an argument to be made, looking at AISI results, that with uncapped tokens it would be just slightly behind the closed weight players.

Re: Kimi-K3 on HuggingFace

#440
post #341

This has to be one of the craziest uploads on the internet up until now. Raw fucking intelligence at your disposal, free to download. If you'd describe what's happening now to someone from five years ago they'd think you're hallucinating or mad.

I think we felt the same when Apache or MySql was releasing back in ancient times.

Apache and Mysql aren't general purpose brains
Post reply on HN