Live data from Hacker News

AMD acquires Taalas to boost inference performance by etching models in silicon

theregister.com

31–40 of 710 posts

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#32
post #15

Earlier quoted context omitted.

If we had deepseek v4 flash 0731 etched on a chip it would be more than capable enough and fast enough for so many people's needs, even hardcore engineer.

Will be capable and fast enough for 2-3 weeks until new sota drops

If it is capable today why would a new model change this?

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#33
post #17

Earlier quoted context omitted.

It doesn’t believe it’s running on that chip, it’s arguing with me

It's running a very small, non-reasoning model at the moment. But more generally, almost all LLMs argue on the hardware/model they are/are on.

Which model? Or how many active parameters?

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#35

Earlier quoted context omitted.

It's running a very small, non-reasoning model at the moment. But more generally, almost all LLMs argue on the hardware/model they are/are on.

Which model? Or how many active parameters?

Llama 3.1 8B model

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#37

Honestly, this is starting to make more and more sense. SOTA models are starting to converge to certain architecture and capabilities. I wouldn’t be surprised we end up with a base model ASIC + “fine tune” card where it’s a physical LoRA style adapter.

Wouldn't this mean someone with sufficient hardware could lift the SOTA model weights off the chip? Or are you saying that these chips would only be used internally by these companies and not sold to the public?

I wouldn't expect companies not sharing their weights today to be any more likely to share them if they're on hardware, this doesn't sufficiently hide weights from a local user.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#38

AMD could have saved their money and used their own hardware! I've got a language model doing 60k tok/s on AMD hardware already, a Xilinx Kria K26 SOM, with the weights baked into URAM/BRAM with zero DRAM in the token loop. Same thesis as Taalas: single-stream decode is bandwidth bound, so stop fetching weights from far away. Caveats stacked high, obviously. It's 3.16M parameters (tinystories, and I also have a kevin…

Well, technically it is their hardware now...

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#40

Honestly, this is starting to make more and more sense. SOTA models are starting to converge to certain architecture and capabilities. I wouldn’t be surprised we end up with a base model ASIC + “fine tune” card where it’s a physical LoRA style adapter.

Then we can have machine psychologists pull cards when they run amok.
Post reply on HN