> At 20 billion parameters per chip, you’d need just 50 accelerators to support a trillion-parameter model I don't see any evidence that this is possible. From my understanding, the whole model needs to be on a single chip. Which rules out any popular frontier models with several trillions of parameters. Even smaller sub-frontier models have hundreds of millions of parameters, so these would be ruled out as well.
AMD acquires Taalas to boost inference performance by etching models in silicon
151–160 of 712 posts
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#152"You are not prepared" --Illidan Stormrage
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#153Honestly, this is starting to make more and more sense. SOTA models are starting to converge to certain architecture and capabilities. I wouldn’t be surprised we end up with a base model ASIC + “fine tune” card where it’s a physical LoRA style adapter.
Then we can have machine psychologists pull cards when they run amok.
Now your robot can respond sarcastically when you ask for chicken nuggets. Again. It also doesn't dent your walls anymore.
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#154I'm surprised neither OpenAI nor Anthropic made this move first. The Chinese open weight models are pulling ahead and commoditizing their value proposition. Baking models onto silicon would've been the next logical move to get a moat. Google is already doing this and has an experimental project on top of already having TPUs and cramming their quantized flash onto individual TPUs for inference.
A model can't be updated, and a chip that is only relevant for 6 months at max?
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#155Earlier quoted context omitted.
Because Openai and anthropic are not hardware companies. They outsource that to Broadcom and AWS' Annapurna labs.
OpenAI and Anthropic are both designing ASICs.
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#156AMD could have saved their money and used their own hardware! I've got a language model doing 60k tok/s on AMD hardware already, a Xilinx Kria K26 SOM, with the weights baked into URAM/BRAM with zero DRAM in the token loop. Same thesis as Taalas: single-stream decode is bandwidth bound, so stop fetching weights from far away. Caveats stacked high, obviously. It's 3.16M parameters (tinystories, and I also have a kevin…
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#157Given the fast churn of the models, how does it work out? Won’t the silicon etched model already be 1 or more versions behind by the time the silicon comes out. Though if it’s cheap enough, there certainly can be a market for cheaper model inferences.
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#158The demo: https://chatjimmy.ai/
....damn. It's very impressive notwithstanding its limitations.
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#159This is neat but IMO a little crazy. Something I personally haven’t seen much of, in all the discussions of model benchmarks and AI breakthroughs, is a distinction between “peak performance” and “reliable performance”. The “peak performance” of frontier models is very high: they’re solving open math problems, analyzing large codebases, etc. But my subjective impression is that “reliable performance” is mid at best: o…
What are some examples?
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#160I'm surprised neither OpenAI nor Anthropic made this move first. The Chinese open weight models are pulling ahead and commoditizing their value proposition. Baking models onto silicon would've been the next logical move to get a moat. Google is already doing this and has an experimental project on top of already having TPUs and cramming their quantized flash onto individual TPUs for inference.
Apparently Anthropic is moving that way: https://arstechnica.com/ai/2026/08/anthropic-confirms-plans-...
The Jalapeño mentioned («Anthropic is not alone in walking this path») in the article is still a classical Von Neumann architecture.
And Taalas' idea makes sense in a perspective of scale - producing a large number of cards; "for internal use" (a lower order of items) means a high production cost.