Live data from Hacker News

Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

emergingtrajectories.com

331–340 of 349 posts

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#331

The open weight, open architecture releases of the past several days has me more convinced that ultimately, the winner will be whoever burns their models to ASICs fastest. The LLMs themselves are capable of doing some aspects of chip design as evinced by the K3 press release. Furthermore, the frontier models are "good enough" for a wide swathe of tasks and will soon hit that threshold for a good amount of software en…

I'm not convinced, mostly because things like crypto, which I believe went into ASICs, were based on very slowly moving and mostly understood algorithms. LLMs and model architectures seems significantly more volatile. I wouldn't want to be working out the finer details of my chip rollout only to find a new paper/approach that give multiples of performance. So I guess it depends on how much the latest-greatest model m…

I think it depends on how good is "good enough". Personally, I think as long as tasks are not coding, there would be a ton of "good enough" value in having your own personal Google on a hardware device. I think the company that designs and licenses the architecture that packages the silicon for the largest market share will be the winner in commercializing this product. Wouldn't be surprised if this is part of Apple's roadmap to retaining existing customer loyalty.

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#332
post #277

Earlier quoted context omitted.

No the weights are in the metal layers, they cannot be updated.

The base weights can't be updated but from what I recall it allows adding a low rank adapter to customize the model a little bit.

Yes that's what I've read. As far as I know the approach should transfer well to hybrid model architectures like modern Qwen and sizes like 27B by using multiple chips. LoRA-steered Qwen 27B at 10K+ tokens per second would be transformative for some workflows.

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#333

Earlier quoted context omitted.

that stuff in particular -- AI generated with heavy heavy prompting and up front design work and post-implementation testing I have CUDA work here somewhere too but I have the repository private right now

Very cool. Thanks for sharing.

Here's what is key: never rely fully on the "intelligence" of the model. Build a workflow and/or tooling that allows for empirical improvement. e.g. for performance tuning I have a benchmarking framework and a container/server that contains the differential results available via MCP for the agent to observe as it works. Using /goal and a clear destination you want to get to, it will literally grind for hours and use that help. Doesn't help with architectural stuff, but it does help with producing efficient code.

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#334
post #5

I think the risk is overstated. For one, on the margin people are willing to pay a lot for slightly better models. I know personally the value the LLM adds to my workflow is considerably more than the $200/m I pay the frontier labs. I have no interest in optimizing that to get it slightly lower. There are a very vocal minority that optimizes this or companies whose LLM expense is marginal, but I think that's the mino…

> I know personally the value the LLM adds to my workflow is considerably more than the $200/m I pay the frontier labs.

AFAIK the frontier labs make their money on enterprise, and they recently changed their terms so that most businesses can't use the $200/m plan anymore.

This is why companies like Uber - and mine :( - announced per employee AI budgets on the order of ~$1000/month.

When you're forced to pay API costs, the value prop of Kimi/Qwen becomes a lot more compelling.

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#335

Earlier quoted context omitted.

Kimi K3 is still worse than Fable and Fable was trained >4 months ago.

Why are you using the product release date for Kimi K3 and the training date for Fable? Either use the release date for both (6 weeks apart) or if you have it the training date for both.

because its a matter of public record that the release of Fable was artificially delayed

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#336

Earlier quoted context omitted.

> Meaning that in a couple of years you may There is no "may" here. You will see this. It's always difficult to see it from the present, but we're not at some end stage in hardware development; we're still on the same curve our predecessors also couldn't see: they couldn't imagine that there would be high performance computers carried in our pockets, with staggering amounts of storage and compute, putting to shame th…

Oh yeah the may is on the 2 year time horizon. It could be 3 or 4. Or next year.

Yes, the timelines are highly uncertain, but the AI boom and AI money are reinvigorating the entire fabrication ecosystem from top to bottom and causing innovation where there had only been incremental progress for years. People aren't just thinking about their node schedule and yields: they're pursuing radical new designs with far greater urgency.

That's where OpenAI in a box is going to come from. And it won't take long.

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#337

Earlier quoted context omitted.

Agreed 100%. This guy thinks there's a limit on the demand for intelligence. You think that Fable 7 which can run a billion dollar corporation on its own has no consumer demand just because we have fable 5 at 9k tok/s? Who do you think will be the biggest customer of such a model? Fable 7, obviously.

theoretically theres a no limit on the demand of anything if the price is right pretty stupid statement lmao

Marco econ 101 disagrees :) Ask your favorite LLM to explain the "yield curve" to you

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#338
post #134

The open weight, open architecture releases of the past several days has me more convinced that ultimately, the winner will be whoever burns their models to ASICs fastest. The LLMs themselves are capable of doing some aspects of chip design as evinced by the K3 press release. Furthermore, the frontier models are "good enough" for a wide swathe of tasks and will soon hit that threshold for a good amount of software en…

A Fable 5 model running at 9,000 tokens/s on an ASIC rather than 150 tokens/s on electricity chugging Nvidia GPUs, or even giant SRAM Cerebras or Groq chips could be good enough to meet the majority of demand. 640K ought to be enough for anybody.

> 640K ought to be enough for anybody.

The question is really whether 640k is enough to last you until your next hardware upgrade.

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#339

Earlier quoted context omitted.

Why are you using the product release date for Kimi K3 and the training date for Fable? Either use the release date for both (6 weeks apart) or if you have it the training date for both.

because its a matter of public record that the release of Fable was artificially delayed

By Anthropic. The government blocked it after it was already released.

Every LLM product goes through testing and alignment after training, maybe even some quick improvements here and there. Kimi probably did something similar.

Put another way, if Google says they have the best model in the world but won’t release it in December I will start caring in December, not before.

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#340

Earlier quoted context omitted.

> even though I have faster models available I still go to Fable or GPT-5.6 90% of the time What about all the things you don't currently use an LLM for? If a specialized chip can run a model 100 times faster, you can suddenly use it for a lot of things at sub-second latency. You can write "make white transparent and add a red outline to x.png" instead of the corresponding imagemagick invocation and perceive little t…

or, hear me out, advertisers can do real-time advertising based on hyper-now context

Yeah... savage capitalism will ruin this too :'(
Post reply on HN