Live data from Hacker News

AMD acquires Taalas to boost inference performance by etching models in silicon

theregister.com

201–210 of 712 posts

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#201
I've had an endgame idea in mind for a while.

Models, probably first open weight ones like Kimi K3 class, are etched into silicon like this and sold as cartridges almost like old school game cartridges.

You buy a USB-C dongle that the cartridge goes into, or for data centers you have PCI cards that take these in slots.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#202

This is neat but IMO a little crazy. Something I personally haven’t seen much of, in all the discussions of model benchmarks and AI breakthroughs, is a distinction between “peak performance” and “reliable performance”. The “peak performance” of frontier models is very high: they’re solving open math problems, analyzing large codebases, etc. But my subjective impression is that “reliable performance” is mid at best: o…

> out of 100 random questions I might think to ask, it’s likely to say something wrong or stupid a handful of times at least. What are some examples?

There's a benchmark for this and a lot of models get negative scores because they're so unreliable: https://artificialanalysis.ai/evaluations/omniscience

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#203

Earlier quoted context omitted.

A model can't be updated, and a chip that is only relevant for 6 months at max?

People already buy new phones every year, this just creates even more reason to do so

Your location/income bias is showing. Most people do not buy new phones every year.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#205
post #201

I've had an endgame idea in mind for a while. Models, probably first open weight ones like Kimi K3 class, are etched into silicon like this and sold as cartridges almost like old school game cartridges. You buy a USB-C dongle that the cartridge goes into, or for data centers you have PCI cards that take these in slots.

Each cartridge costs $1,000. Do you still want it?

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#206
post #31

The demo: https://chatjimmy.ai/

I freakin' love this demo. It feels magical.

For those old enough to remember, this is like dial up internet to broadband. So fast it creates new markets

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#207
post #121

Given the fast churn of the models, how does it work out? Won’t the silicon etched model already be 1 or more versions behind by the time the silicon comes out. Though if it’s cheap enough, there certainly can be a market for cheaper model inferences.

Already models have gotten really good at a lot of things.

A lot of people would probably be happy to stick with the same model for a year or two if it’s 10x faster and cheaper.

And perhaps older models can become cheaper over time as newer models come out on new silicon for a higher price. That incentivizes people to stick with older models.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#208

AMD could have saved their money and used their own hardware! I've got a language model doing 60k tok/s on AMD hardware already, a Xilinx Kria K26 SOM, with the weights baked into URAM/BRAM with zero DRAM in the token loop. Same thesis as Taalas: single-stream decode is bandwidth bound, so stop fetching weights from far away. Caveats stacked high, obviously. It's 3.16M parameters (tinystories, and I also have a kevin…

Yeah Im surprised nobody is talking about this. When everyone first saw Taalas I looked at the design and it had a big legup in physical cache availale compared to most chips. Makes you wonder how much of a benefit there is to the actual "baking" of the model vs just having a large chip with a ton of SRAM (or whatever) soldered close to the edge physically. I feel like what we really need is the ability to solder com…

Taalas does not have cache so...

I agree that Groq with multilayer hybrid bonding could be a good idea.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#209

I'm surprised neither OpenAI nor Anthropic made this move first. The Chinese open weight models are pulling ahead and commoditizing their value proposition. Baking models onto silicon would've been the next logical move to get a moat. Google is already doing this and has an experimental project on top of already having TPUs and cramming their quantized flash onto individual TPUs for inference.

Personally I think Apple should have acquired them. if you could burn a gemma4 class model into an iphone and actually get extremely low latency and low battery usage it would feel like the future IMO. even if it means you wont get frontier intelligence, there might actually be incentive to buy a new mobile device every year again.

Slightly besides your point, but it's interesting how many here naturally ponder about how the current winner could or "should" keep winning, instead of how another company could become a competitor by doing the more clever thing the incumbent isn't thinking about.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#210

The demo: https://chatjimmy.ai/

The speed is awesome, in the true sense of the word. It's great at knowledge and basic stuff but the output is complete junk for anything concerning new facts or slightly esoteric topics.
Post reply on HN