Live data from Hacker News

Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

emergingtrajectories.com

191–200 of 349 posts

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#191

The open weight, open architecture releases of the past several days has me more convinced that ultimately, the winner will be whoever burns their models to ASICs fastest. The LLMs themselves are capable of doing some aspects of chip design as evinced by the K3 press release. Furthermore, the frontier models are "good enough" for a wide swathe of tasks and will soon hit that threshold for a good amount of software en…

Another reason this may be true, selling hardware may be the only durable moat after a while

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#192

The open weight, open architecture releases of the past several days has me more convinced that ultimately, the winner will be whoever burns their models to ASICs fastest. The LLMs themselves are capable of doing some aspects of chip design as evinced by the K3 press release. Furthermore, the frontier models are "good enough" for a wide swathe of tasks and will soon hit that threshold for a good amount of software en…

> the winner will be whoever burns their models to ASICs fastest.

It'll obviously be China, and they won't need the bleeding edge of lithography tech to make it happen. Every Chinese smartphone will have something like Sonnet 5, along your car (well, not those of us in the US, but we'll look longingly at pictures of them while we drive whatever the government decides we're allowed to drive in Fortress America).

Give it ten years and your smart litterbox from Temu will be running its own local model.

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#193

Earlier quoted context omitted.

> the winner will be whoever burns their models to ASICs fastest. I don't understand how this works when the models are evolving so fast that your burned ASIC is outdated (or at least not top of the line) in a few weeks.

If llama 3 70b were available for $400 today, people would make it work. They would have 4 of them working side by side or 1 of them working with a GPU on another model. To me, it's like imagining if Sonnet 3 was burned into an ASIC 8 years ago and then never changed. It would still be revolutionary, and today we would have an entire ecosystem of tools and services built around it, likely surpassing some of our curre…

I think you'd be better off getting 1% of a $40,000 card. Or maybe 10% of a $4,000 card is more realistic.

I think that's the problem at the moment. A much better LLM that doesn't use my battery is 20ms and <10mb of data away.

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#194
post #29

So, what's the most affordable way for a pleb who doesn't own 17 H100s to use Kimi K3 or Qwen 3.8?

I haven't seen either of these running outside their creator's services yet, but typically you can watch services like openrouter or nano-gpt for it to show up at a (usually small) discount.

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#195

Earlier quoted context omitted.

Interesting. Is the speedup from specializing for the shape of Llama 3.1 or are they (contra my mental model) actually winning on burning in the weights?

The weights are in SRAM, so the LLM architecture is burned in but the weights can be updated.

No the weights are in the metal layers, they cannot be updated.

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#196

Earlier quoted context omitted.

No, if you burn your model into the silicon you don't need memory as the output of a layer flows through circuitry directly into the next one. No I/O to memory of any kind. You still need a bit of memory for in flight answers, but that's it.

You can pipeline in software too, and even with a dedicated circuit, you need to put weights somewhere , because you're not going to have a special partially-applied FP8 FMA unit for each of your trillion model weights. I can accept the idea of specializing a circuit for a specific model shape , but I'm not seeing a need to specialize a circuit for the weights inside the shape.

Apparently that is what Taalas did though. Not an hardware person, so take this with a pinch of salt

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#197

The open weight, open architecture releases of the past several days has me more convinced that ultimately, the winner will be whoever burns their models to ASICs fastest. The LLMs themselves are capable of doing some aspects of chip design as evinced by the K3 press release. Furthermore, the frontier models are "good enough" for a wide swathe of tasks and will soon hit that threshold for a good amount of software en…

Apple's M7 comes to mind...

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#199

Earlier quoted context omitted.

Honestly, whether you think burning current SOTA to hardware is an overinvestment risk depends on what your definition of intelligence is. If you think intelligence is something that can grow like height such that 18 months from now we will basically be bowing down to machine god giants that are running on B200s, then investing in ASICs is the wrong move. However, if you you subscribe to the (very reasonable view) th…

At a 50-100x speedup even a GPT-4o class model could perhaps compete with much newer models simply by thinking deeper, doing harness-controlled Ralph loops, etc. Sure, then it might be "only" ~2-5x faster, but, you wouldn't need to throw all the ASICs into the trash bin. One could also imagine hybrid models, where part of the model is burned into ASICs and part of the model exists in VRAM/HBM2 so it can be updated. I…

Giving a literal monkey the ability to press more keys faster doesn't get a good joke from it.

A model will often come up with worse results given more cycles of compute, only because it will tailspin from second guesses, rethinking and literal flip-flopping on concepts.

--- edit: to those following the thread below... if you look at the comments from the account replying, it's pretty obviously a pro-China account and all replies are antagonistic against anything other than a total submission to the Chinese state. My responses are intentionally antagonistic as every point I've brought up is completely ignored in favor of insults, so yeah, I've been insulting back.

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#200

Earlier quoted context omitted.

The tokens per second performance numbers coming from Cerebrus/Talas are several orders of magnitude higher than models running on GPUs, which is such a huge step change that it will enable many more uses of LLMs that are impractical otherwise. I.e. think about gamers and burning in an LLM chip on a game console like a future Play Station - it doesn't matter if its a frontier LLM if it allows them to talk to in game…

I admit I would like a faster model - but even though I have faster models available I still go to Fable or GPT-5.6 90% of the time. So there is a gap between a potential preference and a revealed preference. Custom AI for things like facial recognition in cameras has existed for decades, before LLMs were a thing. I don't see that getting replaced. And on-device conversational intelligence might go that route as well…

What you think doesn't matter unless its strictly for personal use.

Your company will decide what makes economical sense.

Post reply on HN