Live data from Hacker News

Kimi K3: Open Frontier Intelligence

kimi.com

431–440 of 1001 posts

Re: Kimi K3: Open Frontier Intelligence

#431
I didn’t realize that GPT-5.6 is basically dominating the cost/intelligence Pareto frontier right now, at least for this set of benchmarks. Otherwise it’s only Fable on the very high end and DeepSeek on the very low end. This Kimi model gets close, though.

Re: Kimi K3: Open Frontier Intelligence

#432
post #415

> Chip Design > As an early proof of concept, Kimi K3 designed a chip to serve a nano model built on its own architecture. In a single 48-hour autonomous run, K3 built, optimized, and verified the chip using open-source EDA tools on the Nangate 45nm library. Within 4 mm², the chip closes timing at 100 MHz and sustains over 8,700 tokens/s decode throughput in simulation, packing 1.46M standard cells, 0.277 MB of SRAM,…

I had a thought a while back: sell large local models burned onto fused compute / ROM chips. Like cartridges for old game consoles. Slot (or probably plug into USB-C) and go. It’s an ASIC with the model wired into it so it’s very low power and fast. I’d buy these. Say $100 for a frontier class model. Maybe more.

Kinda already exists.

demo https://chatjimmy.ai/

https://news.ycombinator.com/item?id=47103661

Re: Kimi K3: Open Frontier Intelligence

#433
post #415

> Chip Design > As an early proof of concept, Kimi K3 designed a chip to serve a nano model built on its own architecture. In a single 48-hour autonomous run, K3 built, optimized, and verified the chip using open-source EDA tools on the Nangate 45nm library. Within 4 mm², the chip closes timing at 100 MHz and sustains over 8,700 tokens/s decode throughput in simulation, packing 1.46M standard cells, 0.277 MB of SRAM,…

I had a thought a while back: sell large local models burned onto fused compute / ROM chips. Like cartridges for old game consoles. Slot (or probably plug into USB-C) and go. It’s an ASIC with the model wired into it so it’s very low power and fast. I’d buy these. Say $100 for a frontier class model. Maybe more.

You need terabytes of memory to run a frontier class model

Re: Kimi K3: Open Frontier Intelligence

#434
post #415

Earlier quoted context omitted.

I had a thought a while back: sell large local models burned onto fused compute / ROM chips. Like cartridges for old game consoles. Slot (or probably plug into USB-C) and go. It’s an ASIC with the model wired into it so it’s very low power and fast. I’d buy these. Say $100 for a frontier class model. Maybe more.

This would be very compelling. Can anyone share more details on how it would work? Only issue is that you are stuck at a certain point in time but that’s not a huge deal. Even just a good 27b model would be useful.

Talaas have done this with a llama 3 model. Runs at like, 16k/tokens a second oror something obscene. Very little power draw too.

Doesn’t need hbm or lots of memory, because the hardware can just forward the data straight to the next layer and you don’t need to round trip through memory.

They claim to be working on an approach to make the underlying hardware a bit more reusable between models.

Re: Kimi K3: Open Frontier Intelligence

#435
post #415

> Chip Design > As an early proof of concept, Kimi K3 designed a chip to serve a nano model built on its own architecture. In a single 48-hour autonomous run, K3 built, optimized, and verified the chip using open-source EDA tools on the Nangate 45nm library. Within 4 mm², the chip closes timing at 100 MHz and sustains over 8,700 tokens/s decode throughput in simulation, packing 1.46M standard cells, 0.277 MB of SRAM,…

I had a thought a while back: sell large local models burned onto fused compute / ROM chips. Like cartridges for old game consoles. Slot (or probably plug into USB-C) and go. It’s an ASIC with the model wired into it so it’s very low power and fast. I’d buy these. Say $100 for a frontier class model. Maybe more.

I love this for the popular sci-fi trope too, where you see some ship engineer swap one glowing crystal "compute core" for another.

We could have the photonic AI model ASICs for real!

Re: Kimi K3: Open Frontier Intelligence

#436

Earlier quoted context omitted.

The blog post says it's going to be open, but I don't think the weights have been released yet: > Kimi K3 is the first open model to reach 2.8 trillion parameters. It marks the latest step in Kimi's sustained push at the scaling frontier: for nine of the past twelve months, Kimi models have set the upper bound of open-model sizes. https://www.kimi.com/blog/kimi-k3

> The full model weights will be released by July 27, 2026. Still sensible to mark proprietary for now though.

not much reason to think this won't happen except unconfirmed gossip, but I fully expect the next one to not be released. actually I won't be surprised if even this release was withheld and the announcement withdrawn.

Re: Kimi K3: Open Frontier Intelligence

#437
post #415

> Chip Design > As an early proof of concept, Kimi K3 designed a chip to serve a nano model built on its own architecture. In a single 48-hour autonomous run, K3 built, optimized, and verified the chip using open-source EDA tools on the Nangate 45nm library. Within 4 mm², the chip closes timing at 100 MHz and sustains over 8,700 tokens/s decode throughput in simulation, packing 1.46M standard cells, 0.277 MB of SRAM,…

I had a thought a while back: sell large local models burned onto fused compute / ROM chips. Like cartridges for old game consoles. Slot (or probably plug into USB-C) and go. It’s an ASIC with the model wired into it so it’s very low power and fast. I’d buy these. Say $100 for a frontier class model. Maybe more.

https://chatjimmy.ai

Re: Kimi K3: Open Frontier Intelligence

#438

Looks like open models being months behind is a thing of the past. Now more like weeks.

Fable 5 is constrained Mythos, which came out before April

Sol came out (public access restrictions that Chinese models don’t have to worry about) just a week ago.

Re: Kimi K3: Open Frontier Intelligence

#440
According to artificialanalysis, cost per task is $0.94, which is almost the same as $1.04 of gpt 5.6 sol max (fable is most expensive by far, at $2.75). Things like glm 5.2 max cost roughly half that. The model certainly sounds extremely impressive for something not from openai/antrophic, but the price makes it a mediocre product.

Instruction following seems lower than I’d like, too. OTOH scores on agentic stuff seem high, which… feels a bit contradictory? I thought decent instruction following is step 1 of solid agentic workflow.

The benchmarks look nothing short of incredible. Assuming it’s not benchmaxxed to hell and back it’s just a notch below gpt 5.6, which came out what, a week ago? If the performance claims hold up the delayed Gemini 3.5 pro will likely end up not only behind fable, but also behind 5.6 and a (supposed) open weights model. Google might have to do some real soul-searching.

Post reply on HN