Earlier quoted context omitted.
Could we not just make bigger wafers, if the technology called for it?
The investment in bigger machines at the fab might set you back billions. I don't know about the lithography technology either, how easy you can scale it to larger wafers?
AMD acquires Taalas to boost inference performance by etching models in silicon
411–420 of 712 posts
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#412Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#413Given the fast churn of the models, how does it work out? Won’t the silicon etched model already be 1 or more versions behind by the time the silicon comes out. Though if it’s cheap enough, there certainly can be a market for cheaper model inferences.
I think the real value here is not as a customer-facing agent/chatbot but for for automated processes. Think of all the companies out there that have LLMs doing simple tasks like categorizing customer feedback emails. For such tasks, you don't gain much from better models, so if you could run it 10x cheaper on a slightly older model, it would absolutely be worth it. Pretty much any place people are currently running…
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#414Earlier quoted context omitted.
Baking the base models on to ROM makes a lot of economic sense. SRAM for the KV cache & fine-tunes, not so much. Sure you’d get incredible speeds but it’s not scalable from a die-size or cost perspective. Rather base model on ROM + KV cache on DRAM is much more scalable. Also this would work great for edge devices that have a 2-5 year lifecycle.
It is my understanding that just baking the model itself into silicon only gives moderate gains because memory bandwidth remains a bottleneck.
Think it's called Askjimmy or similar.
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#415Earlier quoted context omitted.
Their PoC chips are big, but then it's ridiculously fast (have you seen chatjimmy.ai?). Also they must be holding a bunch of patents.
Its a cool demo, but its gpt-3.5 level stupid, or worse. edit: Ok, I will self-apologize. Its apparently a 3B model. Mighty impressive for what it does.
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#416HN is failing to understand that AMD knows this well.
Taalas has WO2025217724A1 pending and AMD wants that because it is immediately a function block they can sell to anyone doing FP math, since large (mostly) read only memory banks are ideally suited for that micro-code type stuff.
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#417Earlier quoted context omitted.
Could we not just make bigger wafers, if the technology called for it?
The investment in bigger machines at the fab might set you back billions. I don't know about the lithography technology either, how easy you can scale it to larger wafers?
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#418Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#419Earlier quoted context omitted.
Baking the base models on to ROM makes a lot of economic sense. SRAM for the KV cache & fine-tunes, not so much. Sure you’d get incredible speeds but it’s not scalable from a die-size or cost perspective. Rather base model on ROM + KV cache on DRAM is much more scalable. Also this would work great for edge devices that have a 2-5 year lifecycle.
Baking the base models on to ROM makes a lot of economic sense. Less so for consumers though, because it'd mean the phone is out of date in 3 months when a better model comes along.
Ofc if the model has some critical bugs that’s another matter.
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#420Thinking that five or six years from now, Fable-level intelligence could be provided at 100x the current speed... makes me feel lost. I cannot imagine what the future will look like.
It tells me that they have some kind of insider knowledge that the models have hit their limits and won't be getting much better, and it makes sense economically speaking to just bake the current models and use them for the next 5-10 years. Looks like we're near the top of the S curve.
But ”top of the curve reached” feels like the likelier scenario.