Live data from Hacker News

AMD acquires Taalas to boost inference performance by etching models in silicon

theregister.com

131–140 of 712 posts

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#131
post #108

Earlier quoted context omitted.

From what I remember, these chips are not mobile size yet

A small model would be. I think that’s more the point. It’s definitely not SOTA but it’s fast and energy efficient and local.

Nope, a small model would be larger than the whole iPhone SoC.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#132
post #121

Given the fast churn of the models, how does it work out? Won’t the silicon etched model already be 1 or more versions behind by the time the silicon comes out. Though if it’s cheap enough, there certainly can be a market for cheaper model inferences.

I find speed alone would be a game changer for current models. I hardly find any task anymore that the current frontier models can't do with max reasoning after several rounds of feedback (provided sufficient instruction and the right harness). But waiting an hour or more for reasoning to finish is getting really cumbersome. If they could do the same in seconds (and for cheap of course), I'm pretty sure we'd pretty soon see major software companies pop up that are run by a single human.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#133

I'm surprised neither OpenAI nor Anthropic made this move first. The Chinese open weight models are pulling ahead and commoditizing their value proposition. Baking models onto silicon would've been the next logical move to get a moat. Google is already doing this and has an experimental project on top of already having TPUs and cramming their quantized flash onto individual TPUs for inference.

Because Openai and anthropic are not hardware companies. They outsource that to Broadcom and AWS' Annapurna labs.

OpenAI and Anthropic are both designing ASICs.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#134

I'm surprised neither OpenAI nor Anthropic made this move first. The Chinese open weight models are pulling ahead and commoditizing their value proposition. Baking models onto silicon would've been the next logical move to get a moat. Google is already doing this and has an experimental project on top of already having TPUs and cramming their quantized flash onto individual TPUs for inference.

Personally I think Apple should have acquired them. if you could burn a gemma4 class model into an iphone and actually get extremely low latency and low battery usage it would feel like the future IMO. even if it means you wont get frontier intelligence, there might actually be incentive to buy a new mobile device every year again.

The Taalas chips are not physically small. And part of their secret (if you look at the design) is just locating a bunch of memory soldered on the edges ( I belive higher amounts of SRAM ? )

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#135

I'm surprised neither OpenAI nor Anthropic made this move first. The Chinese open weight models are pulling ahead and commoditizing their value proposition. Baking models onto silicon would've been the next logical move to get a moat. Google is already doing this and has an experimental project on top of already having TPUs and cramming their quantized flash onto individual TPUs for inference.

Personally I think Apple should have acquired them. if you could burn a gemma4 class model into an iphone and actually get extremely low latency and low battery usage it would feel like the future IMO. even if it means you wont get frontier intelligence, there might actually be incentive to buy a new mobile device every year again.

[flagged]

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#136
post #60

Earlier quoted context omitted.

Considering the rate of model development and rail hopping, seems like baking models into silicon is speed-running obsolescence.

"seems like baking models into silicon is speed-running obsolescence" Now maybe. When models are flying passenger aircraft, other prerogatives will assert themselves. When a 50TB ROM means you can impulse purchase a ChatGPT 6.3 xhigh that runs on batteries, yet more use cases will be apparent.

Well, 50TB ROM Taalas HC1 style would be apparently a 400000b transistor system through a chip sized 2.5 meters on the side... :)

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#138
post #17

The demo: https://chatjimmy.ai/

It doesn’t believe it’s running on that chip, it’s arguing with me

AIs don't intrinsically know anything about themselves so they often give wrong answers to such questions. This can be fixed by putting info in the system prompt but they may consider it a waste of tokens since most usage doesn't benefit from that information.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#139

Earlier quoted context omitted.

Personally I think Apple should have acquired them. if you could burn a gemma4 class model into an iphone and actually get extremely low latency and low battery usage it would feel like the future IMO. even if it means you wont get frontier intelligence, there might actually be incentive to buy a new mobile device every year again.

I don't think this works out from a cost/silicon perspective. Small models already run pretty well in software (since the weights fit in cache) and big models require silicon area proportional to the size of weights. On a mobile device putting a chip like this is competing directly in BOM and power against a whole lot more l3 cache, and the l3 cache makes everything faster

My question is what changes about LLM use cases when you’re getting 1000 tok/s? Models in silicon might dramatically change how we think about them.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#140
post #108

Earlier quoted context omitted.

A small model would be. I think that’s more the point. It’s definitely not SOTA but it’s fast and energy efficient and local.

> A small model would be [mobile size] A ~30mm side for the HC1 tech for an 8b model (still unclear the planned HC2)?

Is that analogue or are they baking floating points into the silicon?
Post reply on HN