AMD acquires Taalas to boost inference performance by etching models in silicon
301–310 of 712 posts
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#302Given the fast churn of the models, how does it work out? Won’t the silicon etched model already be 1 or more versions behind by the time the silicon comes out. Though if it’s cheap enough, there certainly can be a market for cheaper model inferences.
That’s not going to be true forever. As models mature, we will hit diminishing returns. Major improvements will come annually rather monthly - matching the roughly annual release of new processors. Model ROM’s will likely get integrated into die packages just like DRAM now.
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#303100% local and no leaks.
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#304I'm surprised neither OpenAI nor Anthropic made this move first. The Chinese open weight models are pulling ahead and commoditizing their value proposition. Baking models onto silicon would've been the next logical move to get a moat. Google is already doing this and has an experimental project on top of already having TPUs and cramming their quantized flash onto individual TPUs for inference.
I also think that etching models into ASICs may be a bit too inflexible for what OpenAI and Anthropic want.
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#305Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#306Earlier quoted context omitted.
This is only true for people who are solely focused on performance. There is absolutely a market for acceptable performance combined with predictability.
True, but predictability cuts both ways. We're all used to having to constantly update our browsers and phones to keep up with the security arms race. If a frozen model can't be updated, it will predictably remain vulnerable to any "exploits" or idiosyncratic quirks that people discover over time. Let's say, as somebody suggested in another comment, that you buy 100,000 of these chips and deploy them to run fast-food…
I think it's far more likely to see them used in safety critical applications where you need a capable model that can run on low power and doesn't have multiple layers of operating abstractions between the model and the hardware.
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#307Given the fast churn of the models, how does it work out? Won’t the silicon etched model already be 1 or more versions behind by the time the silicon comes out. Though if it’s cheap enough, there certainly can be a market for cheaper model inferences.
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#308Given the fast churn of the models, how does it work out? Won’t the silicon etched model already be 1 or more versions behind by the time the silicon comes out. Though if it’s cheap enough, there certainly can be a market for cheaper model inferences.
Also I believe there is both a market for extremely fast local inference with current model performance and that such fast inference would unlock unforeseen usecases. Especially as TPS approaches early computer clock cycles and data rates.
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#309I'm surprised neither OpenAI nor Anthropic made this move first. The Chinese open weight models are pulling ahead and commoditizing their value proposition. Baking models onto silicon would've been the next logical move to get a moat. Google is already doing this and has an experimental project on top of already having TPUs and cramming their quantized flash onto individual TPUs for inference.
Considering the rate of model development and rail hopping, seems like baking models into silicon is speed-running obsolescence.
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#310Is there any LLM from exactly one year ago that would be worth running? In Aug 2025 you had - OpenAI o3 - Opus 4.1 - Gemini 2.5 Pro - Grok 4 Even if those were almost free to run, you'd be way better off with Deepseek flash 0731 or GPT 5.6 Luna, which already are almost free. Other than for things where the t/s are critical, it seems like a bad idea to etch a model into silicon.