Live data from Hacker News

AMD acquires Taalas to boost inference performance by etching models in silicon

theregister.com

421–430 of 723 posts

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#421

Thinking that five or six years from now, Fable-level intelligence could be provided at 100x the current speed... makes me feel lost. I cannot imagine what the future will look like.

It's not just speed, it will consume a lot less energy per token, maybe even more than 100x difference. And cost for a chip that runs that one model will also go down a lot once volume scales up. They will end up way cheaper than flexible GPU chips.

I expect AI models chopped up into building blocks where 99.9% of the compute is fixed but glued together with flexible "fine tuning" layers that will adapt them to specific applications. Those kind of chips will run 99% of consumer AI and at some point be integrated into consumer devices.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#422

Earlier quoted context omitted.

It is my understanding that just baking the model itself into silicon only gives moderate gains because memory bandwidth remains a bottleneck.

Did you try using the the talaas chat? Something stupid like 18k tokens/second. Think it's called Askjimmy or similar.

Oh boy, thanks for sharing this, truly mind blowing. It was chatjimmy.ai

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#423

I'm surprised neither OpenAI nor Anthropic made this move first. The Chinese open weight models are pulling ahead and commoditizing their value proposition. Baking models onto silicon would've been the next logical move to get a moat. Google is already doing this and has an experimental project on top of already having TPUs and cramming their quantized flash onto individual TPUs for inference.

Personally I think Apple should have acquired them. if you could burn a gemma4 class model into an iphone and actually get extremely low latency and low battery usage it would feel like the future IMO. even if it means you wont get frontier intelligence, there might actually be incentive to buy a new mobile device every year again.

It would be cool if the future was a standard fairphone like module system where you could replace the model chip when you felt like it without having to shell out 1-2k $$$s for a new phone

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#424

The demo: https://chatjimmy.ai/

try let it to get a brief of france history which being reading a while hit the button and then the brieft jump into my eye Generated in 0.051s • 14,092 tok/s Impressive... Given gpt 5.5 was very good to me and gpt 5.6 series seems not boost too much, i kinda like the way bake the model weight to the chip, and connect multiple chip to serve the large scale model and allow respin some parts(ROM like?) to do model weig…

> try let it to get a brief of france history which being reading a while hit the button and then the brieft jump into my eye

WTF is this?

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#425

Thinking that five or six years from now, Fable-level intelligence could be provided at 100x the current speed... makes me feel lost. I cannot imagine what the future will look like.

It makes me dizzy. I have no idea what is going to happen within even a year from now, can barely even imagine it.

I'm trying to get the most out of it by redlining my AI subscriptions. Hopefully I'll manage to start a business in my niche. I don't even know if my niche will exist in the future.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#426

Earlier quoted context omitted.

The Taalas chips are not physically small. And part of their secret (if you look at the design) is just locating a bunch of memory soldered on the edges ( I belive higher amounts of SRAM ? )

Baking the base models on to ROM makes a lot of economic sense. SRAM for the KV cache & fine-tunes, not so much. Sure you’d get incredible speeds but it’s not scalable from a die-size or cost perspective. Rather base model on ROM + KV cache on DRAM is much more scalable. Also this would work great for edge devices that have a 2-5 year lifecycle.

We also have ReRAM (Analog Computing), which also holds a promising future given its efficiency and low power. Though ReRAM of larger size is still a research area.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#427
post #394

Earlier quoted context omitted.

Massive economic simulations with thousands if not millions of agents to front run the global economy and stock market. Fully interactive realtime NPCs in videogames at scale. Recommender systems that simulate individual consumers. Crazy shit

About your first example, isn’t the butterfly effect preventing this from being useful? One agent in your simulation decides to sell, and starts an avalanche, that won’t happen in reality?

When you run tens of thousands of simulations for complex economic models, you actually do want to see the extreme outliers too. I can't recall who said it, but in finance the interconnected incentives make so-called Black Swan events much more likely and frequent than models or theories can comfortably account for.

In a way... when it's finance, they should be maybe called Gray'ish Swans?

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#428

I'm surprised neither OpenAI nor Anthropic made this move first. The Chinese open weight models are pulling ahead and commoditizing their value proposition. Baking models onto silicon would've been the next logical move to get a moat. Google is already doing this and has an experimental project on top of already having TPUs and cramming their quantized flash onto individual TPUs for inference.

Chinese already start making DUV which can do the lower end 7nm. They are winning. Once that 7nm and up market cornered by Chinese, AMD Intel and TSMC and Samsung will have to burn thru bleeding edge depreciation faster perhaps from 7yr down to just 18mths. The CPU they generated will be incredibly expensive. Meanwhile Chinese just keep minting the AI cheaply and more efficiently and inching upwards towards 1.4nm.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#429

The demo: https://chatjimmy.ai/

I know it's a relatively tiny model, but damn, is that thing fast. It also mostly passes the "schlong" test https://pastes.io/YcxSi8Fp

Which model is it?

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#430
post #234

What I like about this, is that it significantly increases the probability of a sci-fi scenario where you're picking up a hot chip on the black market; rumor has it, Mythos 9 weights baked in...

Black market uncensored heretic Mythos weights...
Post reply on HN