Earlier quoted context omitted.
From what I remember, these chips are not mobile size yet
A small model would be. I think that’s more the point. It’s definitely not SOTA but it’s fast and energy efficient and local.
AMD acquires Taalas to boost inference performance by etching models in silicon
131–140 of 712 posts
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#132Given the fast churn of the models, how does it work out? Won’t the silicon etched model already be 1 or more versions behind by the time the silicon comes out. Though if it’s cheap enough, there certainly can be a market for cheaper model inferences.
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#133I'm surprised neither OpenAI nor Anthropic made this move first. The Chinese open weight models are pulling ahead and commoditizing their value proposition. Baking models onto silicon would've been the next logical move to get a moat. Google is already doing this and has an experimental project on top of already having TPUs and cramming their quantized flash onto individual TPUs for inference.
Because Openai and anthropic are not hardware companies. They outsource that to Broadcom and AWS' Annapurna labs.
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#134I'm surprised neither OpenAI nor Anthropic made this move first. The Chinese open weight models are pulling ahead and commoditizing their value proposition. Baking models onto silicon would've been the next logical move to get a moat. Google is already doing this and has an experimental project on top of already having TPUs and cramming their quantized flash onto individual TPUs for inference.
Personally I think Apple should have acquired them. if you could burn a gemma4 class model into an iphone and actually get extremely low latency and low battery usage it would feel like the future IMO. even if it means you wont get frontier intelligence, there might actually be incentive to buy a new mobile device every year again.
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#135I'm surprised neither OpenAI nor Anthropic made this move first. The Chinese open weight models are pulling ahead and commoditizing their value proposition. Baking models onto silicon would've been the next logical move to get a moat. Google is already doing this and has an experimental project on top of already having TPUs and cramming their quantized flash onto individual TPUs for inference.
Personally I think Apple should have acquired them. if you could burn a gemma4 class model into an iphone and actually get extremely low latency and low battery usage it would feel like the future IMO. even if it means you wont get frontier intelligence, there might actually be incentive to buy a new mobile device every year again.
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#136Earlier quoted context omitted.
Considering the rate of model development and rail hopping, seems like baking models into silicon is speed-running obsolescence.
"seems like baking models into silicon is speed-running obsolescence" Now maybe. When models are flying passenger aircraft, other prerogatives will assert themselves. When a 50TB ROM means you can impulse purchase a ChatGPT 6.3 xhigh that runs on batteries, yet more use cases will be apparent.
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#137Burnt in, it needs a zif socket and easy access in every car, aircraft, a pull out slot in a phone, or it's new era planned obselescence.
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#138The demo: https://chatjimmy.ai/
It doesn’t believe it’s running on that chip, it’s arguing with me
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#139Earlier quoted context omitted.
Personally I think Apple should have acquired them. if you could burn a gemma4 class model into an iphone and actually get extremely low latency and low battery usage it would feel like the future IMO. even if it means you wont get frontier intelligence, there might actually be incentive to buy a new mobile device every year again.
I don't think this works out from a cost/silicon perspective. Small models already run pretty well in software (since the weights fit in cache) and big models require silicon area proportional to the size of weights. On a mobile device putting a chip like this is competing directly in BOM and power against a whole lot more l3 cache, and the l3 cache makes everything faster
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#140Earlier quoted context omitted.
A small model would be. I think that’s more the point. It’s definitely not SOTA but it’s fast and energy efficient and local.
> A small model would be [mobile size] A ~30mm side for the HC1 tech for an 8b model (still unclear the planned HC2)?