Live data from Hacker News

Kimi-K3 on HuggingFace

huggingface.co

301–310 of 588 posts

Re: Kimi-K3 on HuggingFace

#301

Earlier quoted context omitted.

The performance bottleneck is not really so much the number of cores or processing power in each core, but the memory bus bandwidth to/from the CPU. I have an older dual socket xeon server here which is a CPU-only LLM test machine with 256GB of RAM and the actual CPU stress is not much, I can even quantify this by how little it spins up the CPU fans to meet thermal load (the CPUs are operating at nowhere near their 1…

Thats where the threadrippers really excelled. They had the lanes for memmory access. We might soon see the return of dinner plate-sized CPUs with thousands of pins.

The epyc Venice SP7 socket is apparently 9324 pins

https://x.com/tomshardware/status/2066846693778510331

Re: Kimi-K3 on HuggingFace

#302
post #24

Earlier quoted context omitted.

No, 27-7 for the rest of the world. The separator is often the only way to distinguish American notation from ISO, so please use a dash for dd-mm-yy and a forward slash for mm/dd/yy

Have never seen 27-7, as someone in rest-of-world

In my time it was “27/VII 1986” in Russia.

Re: Kimi-K3 on HuggingFace

#303

Earlier quoted context omitted.

> My guess is that someone will be selling tokens at less than electricity + depreciation of GCs soon, since there's a lot of competition and "smaller" data centers have overcapacity? This is speculation, correct me if I'm wrong My guess is they are selling you the tokens, then selling your tokens (data) onto someone else.

I see these conspiratorial arguments all the time and I think people massively overestimate the value of the average users tokens. The problems with frontier models (design taste, ability to solve novel/difficult problems, etc) cannot be solved by throwing more slop from the average user at it. Actually, most of the main deficiencies in current models stem from the fact that their data sets aren’t curated and special…

The prompts contain sensitive personal data.

That would be valuable to advertizers for example.

Re: Kimi-K3 on HuggingFace

#304

This will be interesting for a few reasons. First, depending on where the median pricing settles w/ 3rd party providers will tell us what it costs to serve a 3T model. Since it's going to be mxfp4 native, it'll take ~1.5TB of VRAM to host this, which is juuust at the limit of 8xb200s (but realistically you'll need 16x for context / throughput optimisation). Won't be cheap to host, but at least we should get some rang…

> The latest AISI benchmarks on cybersec place it above glm5.2, but still way way behind SotA closed models.

Sota closed models don't even answer cybersecurity questions lol.

Re: Kimi-K3 on HuggingFace

#306
post #106

Earlier quoted context omitted.

> like 180W-250W TDP > running GLM 5.2 would be cool at like ~100 tokens per second for a single session Your power consumption estimates are off for this generation of GPUs. A 27B dense model gets 50-80 tps on an RTX 6000 using 600 watts.

An AMD R9700 gets 20-50 TPS at ~300 watts on 27B. 100 TPS for the 35B MOE model. And there might be some more optimizations to that as AMD software support gets better with ROCm's latest versions.

Which quants? I get these speed (only 20-30 TPS) at Q4_K_M with MTP for 27B on my framework desktop. I draw sub 130W for the whole machine.

Re: Kimi-K3 on HuggingFace

#307

Can't wait to run this at 0.02 tokens/sec on my CPU so I can get a response just in time for next month.

Unfortunately I don't think modern consumer CPUs are physically capable of addressing enough RAM to even load the model into memory. We'd have to wait for some random person to make an extremely quantized version before we could reach those blazing speeds

Re: Kimi-K3 on HuggingFace

#308
post #269

Earlier quoted context omitted.

Yeah, they thing forever and doubt everything "wait but" for 200k tokens for almost any question.

On the flip side, I really like being able to inspect its reasoning chain thoroughly, as opposed to the "black box" that Anthropic models are now.

Right I was going to say, no way of knowing whether these issues are unique to Chinese models.

Re: Kimi-K3 on HuggingFace

#309
post #63

Earlier quoted context omitted.

> Even if the output is like 5-6 tok/s, that might be usable for some purposes. You'll spend ~100x more on electricity than the API cost to have it run on someone else's GPU at several hundred tokens per second. I think some sort of extreme data privacy requirement is the only situation that justifies this, but the intersection of {needs absolute data privacy, needs to run SOTA model, cannot afford GPUs} is really re…

There are a number of use cases where sending the contents of your context and prompts (and the resulting output) to a 3rd party service is off the table as an option, and people will compromise speed for data sovereignty. And not everyone's electricity is equally expensive, I pay about $0.075 USD per kWh. It would for example cost me about $48 a month of electricity (not counting cost of cooling) to run a quad socke…

That's an unusually low electric rate for the US - way below the lowest state average which is Idaho at 12.4 cents. It's certainly possible that you are getting 7.5 cents including delivery, but I've had friends say that they're "getting 13 cents per kWh" here in Massachusetts, but that's just the supply rate and the delivery is another ~18 cents.

There are parts of states like Grant County Washington that have cheap hydro power, but it's very rare for power to be that cheap in the US. Even if this applies to you, it won't apply to the vast majority of people on here who will have electric rates 2-4x higher.

Average electric rates by region:

    New England            28.1 cents
    Mid Atlantic           25.1 cents
    East North Central     20.8 cents
    West North Central     14.8 cents
    South Atlantic         16.1 cents
    East South Central     15.5 cents
    Mountain               14.6 cents
    Pacific Contiguous     26.1 cents
    Pacific Noncontiguous  42.1 cents
https://www.eia.gov/electricity/monthly/epm_table_grapher.ph...

Re: Kimi-K3 on HuggingFace

#310

They release it before I could even get a kimi subscription because of waitlist? lol hard to believe that I might get a kimi subscription from a third party

I just checked and I have it available in opencode go! I just tested it with one message and confirm it works.

It's not practical to use it though. It counts almost 8x more towards your quota than glm 5.2!

https://opencode.ai/docs/go/#usage-limits

Post reply on HN