Live data from Hacker News

Kimi K3: Open Frontier Intelligence

kimi.com

451–460 of 1001 posts

Re: Kimi K3: Open Frontier Intelligence

#451
post #163
post #137

Earlier quoted context omitted.

That's a great question. I just tried "hi" through the same OpenRouter API and the input token count for that was 86 - and for "hi there" the count was 87. I think there's an 85 token hidden system prompt of some sort.

Try {"messages":[ {"role": "user", "content": "hi"} ]} but also an explicitly empty system message: {"messages":[ {"role": "system", "content": ""} {"role": "user", "content": "hi"} ]} and finally {"messages":[ {"role": "system", "content": "x"} {"role": "user", "content": "hi"} ]} Comparing OpenRouter’s tokensPrompt with nativeTokensPrompt can tell you if it came from the provider

I tried prompting "hi" without my own system prompt and it took 86 input tokens, then I set the system prompt to just the word "french" and it jumped up to 99 input tokens. https://gist.github.com/simonw/629b8d05864d7c13e8625a7c48cec...

Re: Kimi K3: Open Frontier Intelligence

#452
post #379

Earlier quoted context omitted.

[flagged]

> I pretty sure OpenAI and Anthropic are doing the same or worse. No they're not. It would end both companies if they were ever found to be doing that. Their terms are clear - if you use the coding plans they can[0] train in return. Enterprise and API, absolutely not. The argument here is that with the Chinese labs you have zero legal recourse. [0] opt-in, thanks

lmao, wasn't xAI caught doing this recently? moreover at least moonshot is being honest about it.

Re: Kimi K3: Open Frontier Intelligence

#453

> Chip Design > As an early proof of concept, Kimi K3 designed a chip to serve a nano model built on its own architecture. In a single 48-hour autonomous run, K3 built, optimized, and verified the chip using open-source EDA tools on the Nangate 45nm library. Within 4 mm², the chip closes timing at 100 MHz and sustains over 8,700 tokens/s decode throughput in simulation, packing 1.46M standard cells, 0.277 MB of SRAM,…

groq did an ASIC for llama and now for nvidia. Their cloud service is fast.

> NVIDIA Groq 3 LPU Inference Accelerator > The NVIDIA Groq 3 LPU is the next generation of Groq’s innovative language processing unit. Each LPX rack features 256 interconnected LPU accelerators that, together with the NVIDIA Vera Rubin platform, supercharge inference. Each LPU accelerator delivers 500 megabytes (MB) of SRAM, 150 terabytes per second (TB/s) of SRAM bandwidth, and 2.5 TB/s scale-up bandwidth.

https://www.nvidia.com/en-us/data-center/lpx/

Re: Kimi K3: Open Frontier Intelligence

#455
post #415

> Chip Design > As an early proof of concept, Kimi K3 designed a chip to serve a nano model built on its own architecture. In a single 48-hour autonomous run, K3 built, optimized, and verified the chip using open-source EDA tools on the Nangate 45nm library. Within 4 mm², the chip closes timing at 100 MHz and sustains over 8,700 tokens/s decode throughput in simulation, packing 1.46M standard cells, 0.277 MB of SRAM,…

I had a thought a while back: sell large local models burned onto fused compute / ROM chips. Like cartridges for old game consoles. Slot (or probably plug into USB-C) and go. It’s an ASIC with the model wired into it so it’s very low power and fast. I’d buy these. Say $100 for a frontier class model. Maybe more.

Interestingly, you could easily run them from the said old consoles! You'd just need a bit of console code to interface (text input/output) with your fully independent LLM subsystem. Imagine Claude for the NES without Internet?

Re: Kimi K3: Open Frontier Intelligence

#456
post #415

> Chip Design > As an early proof of concept, Kimi K3 designed a chip to serve a nano model built on its own architecture. In a single 48-hour autonomous run, K3 built, optimized, and verified the chip using open-source EDA tools on the Nangate 45nm library. Within 4 mm², the chip closes timing at 100 MHz and sustains over 8,700 tokens/s decode throughput in simulation, packing 1.46M standard cells, 0.277 MB of SRAM,…

I had a thought a while back: sell large local models burned onto fused compute / ROM chips. Like cartridges for old game consoles. Slot (or probably plug into USB-C) and go. It’s an ASIC with the model wired into it so it’s very low power and fast. I’d buy these. Say $100 for a frontier class model. Maybe more.

> I’d buy these. Say $100 for a frontier class model. Maybe more.

Sure you would. Running frontier class models on current hardware costs in the order of tens of thousands of dollars. It is more likely that these custom ASICs will be priced competitively with that, and not with Super Mario Bros.

Oh, and energy consumption will be in the same order.

Re: Kimi K3: Open Frontier Intelligence

#458

Earlier quoted context omitted.

Open Source >>> Closed Source [1] I don't want to cheer against my country, but we've given up on open source. The way Anthropic and OpenAI treat their customers as adversaries is embarrassing. I will cheer for China, for Kimi, and for z.ai until we have something in the same category. [1] I'd even be fine with open weights, fair source, or anything that let us have direct access to the weights. Even if that came wit…

I am with you in the spirit of openweights but I am trying to hard-avoid bringing countries into this. The narrative of US vs China only benefits those who want regulatory capture in the US since attacking China is politically much easier than attacking open-weights, so certain groups like to repeatedly call them 'Chinese models'.

It's much more a rallying cry for open weights funding than it is for regulatory capture.

The argument on our side wins - if America or the West don't do open source, China will. And that means -- with certainty -- that China wins the market.

Every politician and VC should hear that loud and clear.

Re: Kimi K3: Open Frontier Intelligence

#459
post #415

> Chip Design > As an early proof of concept, Kimi K3 designed a chip to serve a nano model built on its own architecture. In a single 48-hour autonomous run, K3 built, optimized, and verified the chip using open-source EDA tools on the Nangate 45nm library. Within 4 mm², the chip closes timing at 100 MHz and sustains over 8,700 tokens/s decode throughput in simulation, packing 1.46M standard cells, 0.277 MB of SRAM,…

I had a thought a while back: sell large local models burned onto fused compute / ROM chips. Like cartridges for old game consoles. Slot (or probably plug into USB-C) and go. It’s an ASIC with the model wired into it so it’s very low power and fast. I’d buy these. Say $100 for a frontier class model. Maybe more.

[dead]

Re: Kimi K3: Open Frontier Intelligence

#460
post #441
post #415

Earlier quoted context omitted.

I had a thought a while back: sell large local models burned onto fused compute / ROM chips. Like cartridges for old game consoles. Slot (or probably plug into USB-C) and go. It’s an ASIC with the model wired into it so it’s very low power and fast. I’d buy these. Say $100 for a frontier class model. Maybe more.

Taalas is developing this, but not for Frontier class models. I hope that if we can least get the easy 80% of work done on that sort of hardware, we can greatly reduce the demand for GPUs, HBM and energy to some extent.

There is an amount of brute forcing that becomes possible at those speeds that I think could even take us beyond 80%. If we could have Qwen3.6-27B running at 15k t/s, run 100 attempts concurrently, select top-K solutions and synthesize a final result from them.

There was a paper a while back that showed top-K selection like that with tiny models was able to reliably solve some 1M-step Tower of Hanoi when no frontier model could. Very big level up in capability just from horizontally scaling compute.

Post reply on HN