Live data from Hacker News

AMD acquires Taalas to boost inference performance by etching models in silicon

theregister.com

171–180 of 712 posts

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#171

Earlier quoted context omitted.

"seems like baking models into silicon is speed-running obsolescence" Now maybe. When models are flying passenger aircraft, other prerogatives will assert themselves. When a 50TB ROM means you can impulse purchase a ChatGPT 6.3 xhigh that runs on batteries, yet more use cases will be apparent.

Well, 50TB ROM Taalas HC1 style would be apparently a 400000b transistor system through a chip sized 2.5 meters on the side... :)

Phones were getting too thin anyways.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#172

Is there any LLM from exactly one year ago that would be worth running? In Aug 2025 you had - OpenAI o3 - Opus 4.1 - Gemini 2.5 Pro - Grok 4 Even if those were almost free to run, you'd be way better off with Deepseek flash 0731 or GPT 5.6 Luna, which already are almost free. Other than for things where the t/s are critical, it seems like a bad idea to etch a model into silicon.

That is fkin wild. o3 was just a year ago? The progress is truly insane.

Yeah I had to double check, o3 feels like it was ages ago. But GPT 5 came out Aug 7, so it's only one day off from my 1 year ago cutoff!

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#173
post #60

Earlier quoted context omitted.

Considering the rate of model development and rail hopping, seems like baking models into silicon is speed-running obsolescence.

I'd gladly pay for a Claude Opus 4.6 Thinking High in silicon and use it for 1-2 years. It's good enough for many coding tasks.

It costs something like $300,000 for the hardware to run a model of that size. You'd pay that for a single model for 1-2 years? Not even the AI companies can justify that kind of spend which is why they keep extending the expected lifespan on their hardware in the accounting.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#174
post #77

Earlier quoted context omitted.

It's a terrible moat. You etch the silicon then nobody wants to run it in 6 months because models have advanced that much further.

OTOH, people get a new iPhone every year and they are ok with it.

How is that in any way related to a consumer device? This method doesn't reduce physical memory requirements, so still results in huge die area. This isn't a for-end-user thing, probably for decades.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#175

Earlier quoted context omitted.

A model can't be updated, and a chip that is only relevant for 6 months at max?

People already buy new phones every year, this just creates even more reason to do so

Outside of this website I've never met a person who buys a new phone every year. It's closer to every 3-4 years for most people.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#177
post #174

Earlier quoted context omitted.

OTOH, people get a new iPhone every year and they are ok with it.

How is that in any way related to a consumer device? This method doesn't reduce physical memory requirements, so still results in huge die area. This isn't a for-end-user thing, probably for decades.

Ok, how long until nvidia gives us a new GPU?

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#178

Earlier quoted context omitted.

I'd gladly pay for a Claude Opus 4.6 Thinking High in silicon and use it for 1-2 years. It's good enough for many coding tasks.

It costs something like $300,000 for the hardware to run a model of that size. You'd pay that for a single model for 1-2 years? Not even the AI companies can justify that kind of spend which is why they keep extending the expected lifespan on their hardware in the accounting.

I'm expecting the Taalas MSIC version to cost a fraction of that. Then probably have some kind of cheap subscription to Anthropic for updates (yes, Taalas chips can receive a certain kind of updates: they have a small SRAM).

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#179

Earlier quoted context omitted.

My question is what changes about LLM use cases when you’re getting 1000 tok/s? Models in silicon might dramatically change how we think about them.

That likely isn't as relevant for on-device iPhone usage as it is for Real Work™. I won't notice the difference between 50tps and 1000tps when asking Siri a question.

I don't know. As others have said, the Taalas chip wasn't small, or particularly low power, so it's hard to "imagine" what that tech in an cell phone chip might look like.

But if the basic premise of "good enough LLM at insane throughput" holds, I think it could qualitatively change local uses of LLMs. At a certain speed point, you're able to move from request -> response to a cascade of tool calling and "subagents", which could allow a small model to be much more useful, if provided with a lot of local data and tool calls.

That said, this is assuming you could stuff a "good enough" model into a phone with Taalas-like technology. The Taalas tech demo was an 8B parameter model and required hundreds of watts (IIRC) to run. The efficiency was good given the speed (as I understand), but it's not clear at all that the approach scales small enough to be a sensible coprocessor on an iPhone or whatever.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#180
post #31

Earlier quoted context omitted.

I freakin' love this demo. It feels magical.

I had the same reaction but then I showed it to my partner. She completely didn't get it, in her words "how can it be thinking of a good answer when it's that quick?" I tried to explain but I fear were probably going to be adding artificial sleeps to these things to convince the masses it's doing something clever.

Wait, is it even thinking? Or is it an instant model?
Post reply on HN