I'm surprised neither OpenAI nor Anthropic made this move first. The Chinese open weight models are pulling ahead and commoditizing their value proposition. Baking models onto silicon would've been the next logical move to get a moat. Google is already doing this and has an experimental project on top of already having TPUs and cramming their quantized flash onto individual TPUs for inference.
AMD acquires Taalas to boost inference performance by etching models in silicon
71–80 of 712 posts
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#72I'm surprised neither OpenAI nor Anthropic made this move first. The Chinese open weight models are pulling ahead and commoditizing their value proposition. Baking models onto silicon would've been the next logical move to get a moat. Google is already doing this and has an experimental project on top of already having TPUs and cramming their quantized flash onto individual TPUs for inference.
Considering the rate of model development and rail hopping, seems like baking models into silicon is speed-running obsolescence.
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#73so qwen3.x-27b on hardware? or better deepseek-v4-flash on hardware .
I wrote them an email asking for PrismML Bonsai 27b Ternary which is like 6b or something crazy small and would be a lot easier for them to do initially.
Bonsai Ternary (1.7bits/weight) is a compromise, compromise that has to make sense in the context - efficient when translated into transistors.
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#74I'm surprised neither OpenAI nor Anthropic made this move first. The Chinese open weight models are pulling ahead and commoditizing their value proposition. Baking models onto silicon would've been the next logical move to get a moat. Google is already doing this and has an experimental project on top of already having TPUs and cramming their quantized flash onto individual TPUs for inference.
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#75Imagine the size of chip needed to 'etch' something like Qwen 3.6 27B in size.
Interesting thought, because it's a yield question. How tolerant are models today to a few broken weights. If tolerant, they could churn out many cheaper chips, some perhaps with slight abnormal tendencies ;)
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#76Can anyone comment on the economics and likely turnaround times of this process, when it’s more mature? Would it be realistic for a frontier lab to deploy this or would the turnaround time mean the model is always too out of date? Assuming the weights and architecture are eventually stable, how much cheaper would this end up being?
There are always uses for outdated models. Claude Code is still using haiku 4.5 from ages ago for explore subagents for instance. Not to mention production uses like customer service that only need to be "good enough"
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#77I'm surprised neither OpenAI nor Anthropic made this move first. The Chinese open weight models are pulling ahead and commoditizing their value proposition. Baking models onto silicon would've been the next logical move to get a moat. Google is already doing this and has an experimental project on top of already having TPUs and cramming their quantized flash onto individual TPUs for inference.
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#78Earlier quoted context omitted.
Considering the rate of model development and rail hopping, seems like baking models into silicon is speed-running obsolescence.
Not sure. You can fix the transistors but leave the connections between them open for flexibility, so you only need to change the manufacturing process for the upper masks for every new model.
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#79The demo: https://chatjimmy.ai/
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#80Earlier quoted context omitted.
Llama 3.1 8B model
So this demo is around 90 times faster than typical speeds for the same model at openrouter, and around 30 times faster than the absolute fastest option available (Groq).