Live data from Hacker News

AMD acquires Taalas to boost inference performance by etching models in silicon

theregister.com

41–50 of 710 posts

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#41

Earlier quoted context omitted.

Will be capable and fast enough for 2-3 weeks until new sota drops

If it is capable today why would a new model change this?

Because new stuff instantly makes anything prior bad and incapable and garbage of course! Did you forget the hype-machine speaking notes??? /s

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#42
post #25

I feel like NAND process tech could become useful at solving some of these problems. A GPU where you can update the weights a few thousand times may be sufficient.

NAND hasn't been scaling great lately. It seems like PCM or MRAM would both be better fits.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#43

Earlier quoted context omitted.

Will be capable and fast enough for 2-3 weeks until new sota drops

If it is capable today why would a new model change this?

if capability is a commodity then the differentiator becomes taste.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#44
I'm surprised neither OpenAI nor Anthropic made this move first. The Chinese open weight models are pulling ahead and commoditizing their value proposition.

Baking models onto silicon would've been the next logical move to get a moat.

Google is already doing this and has an experimental project on top of already having TPUs and cramming their quantized flash onto individual TPUs for inference.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#46
post #17

Earlier quoted context omitted.

It doesn’t believe it’s running on that chip, it’s arguing with me

It's running a very small, non-reasoning model at the moment. But more generally, almost all LLMs argue on the hardware/model they are/are on.

What would tokens/sec performance look like for a reasoning model? An order of magnitude slower?

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#49
post #25

I feel like NAND process tech could become useful at solving some of these problems. A GPU where you can update the weights a few thousand times may be sufficient.

The basis of Taalas is "compute in memory" electronics - past Von Neumann's separation of processor and memory.

You need to be able to add|mul where the data (the weights) are stored.

Post reply on HN