The open weight, open architecture releases of the past several days has me more convinced that ultimately, the winner will be whoever burns their models to ASICs fastest. The LLMs themselves are capable of doing some aspects of chip design as evinced by the K3 press release. Furthermore, the frontier models are "good enough" for a wide swathe of tasks and will soon hit that threshold for a good amount of software en…
I'm not convinced, mostly because things like crypto, which I believe went into ASICs, were based on very slowly moving and mostly understood algorithms. LLMs and model architectures seems significantly more volatile. I wouldn't want to be working out the finer details of my chip rollout only to find a new paper/approach that give multiples of performance. So I guess it depends on how much the latest-greatest model m…
Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling
331–340 of 349 posts
Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling
#332Earlier quoted context omitted.
No the weights are in the metal layers, they cannot be updated.
The base weights can't be updated but from what I recall it allows adding a low rank adapter to customize the model a little bit.
Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling
#333Earlier quoted context omitted.
that stuff in particular -- AI generated with heavy heavy prompting and up front design work and post-implementation testing I have CUDA work here somewhere too but I have the repository private right now
Very cool. Thanks for sharing.
Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling
#334I think the risk is overstated. For one, on the margin people are willing to pay a lot for slightly better models. I know personally the value the LLM adds to my workflow is considerably more than the $200/m I pay the frontier labs. I have no interest in optimizing that to get it slightly lower. There are a very vocal minority that optimizes this or companies whose LLM expense is marginal, but I think that's the mino…
AFAIK the frontier labs make their money on enterprise, and they recently changed their terms so that most businesses can't use the $200/m plan anymore.
This is why companies like Uber - and mine :( - announced per employee AI budgets on the order of ~$1000/month.
When you're forced to pay API costs, the value prop of Kimi/Qwen becomes a lot more compelling.
Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling
#335Earlier quoted context omitted.
Kimi K3 is still worse than Fable and Fable was trained >4 months ago.
Why are you using the product release date for Kimi K3 and the training date for Fable? Either use the release date for both (6 weeks apart) or if you have it the training date for both.
Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling
#336Earlier quoted context omitted.
> Meaning that in a couple of years you may There is no "may" here. You will see this. It's always difficult to see it from the present, but we're not at some end stage in hardware development; we're still on the same curve our predecessors also couldn't see: they couldn't imagine that there would be high performance computers carried in our pockets, with staggering amounts of storage and compute, putting to shame th…
Oh yeah the may is on the 2 year time horizon. It could be 3 or 4. Or next year.
That's where OpenAI in a box is going to come from. And it won't take long.
Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling
#337Earlier quoted context omitted.
Agreed 100%. This guy thinks there's a limit on the demand for intelligence. You think that Fable 7 which can run a billion dollar corporation on its own has no consumer demand just because we have fable 5 at 9k tok/s? Who do you think will be the biggest customer of such a model? Fable 7, obviously.
theoretically theres a no limit on the demand of anything if the price is right pretty stupid statement lmao
Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling
#338The open weight, open architecture releases of the past several days has me more convinced that ultimately, the winner will be whoever burns their models to ASICs fastest. The LLMs themselves are capable of doing some aspects of chip design as evinced by the K3 press release. Furthermore, the frontier models are "good enough" for a wide swathe of tasks and will soon hit that threshold for a good amount of software en…
A Fable 5 model running at 9,000 tokens/s on an ASIC rather than 150 tokens/s on electricity chugging Nvidia GPUs, or even giant SRAM Cerebras or Groq chips could be good enough to meet the majority of demand. 640K ought to be enough for anybody.
The question is really whether 640k is enough to last you until your next hardware upgrade.
Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling
#339Earlier quoted context omitted.
Why are you using the product release date for Kimi K3 and the training date for Fable? Either use the release date for both (6 weeks apart) or if you have it the training date for both.
because its a matter of public record that the release of Fable was artificially delayed
Every LLM product goes through testing and alignment after training, maybe even some quick improvements here and there. Kimi probably did something similar.
Put another way, if Google says they have the best model in the world but won’t release it in December I will start caring in December, not before.
Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling
#340Earlier quoted context omitted.
> even though I have faster models available I still go to Fable or GPT-5.6 90% of the time What about all the things you don't currently use an LLM for? If a specialized chip can run a model 100 times faster, you can suddenly use it for a lot of things at sub-second latency. You can write "make white transparent and add a red outline to x.png" instead of the corresponding imagemagick invocation and perceive little t…
or, hear me out, advertisers can do real-time advertising based on hyper-now context