Earlier quoted context omitted.
"seems like baking models into silicon is speed-running obsolescence" Now maybe. When models are flying passenger aircraft, other prerogatives will assert themselves. When a 50TB ROM means you can impulse purchase a ChatGPT 6.3 xhigh that runs on batteries, yet more use cases will be apparent.
Well, 50TB ROM Taalas HC1 style would be apparently a 400000b transistor system through a chip sized 2.5 meters on the side... :)
AMD acquires Taalas to boost inference performance by etching models in silicon
171–180 of 712 posts
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#172Is there any LLM from exactly one year ago that would be worth running? In Aug 2025 you had - OpenAI o3 - Opus 4.1 - Gemini 2.5 Pro - Grok 4 Even if those were almost free to run, you'd be way better off with Deepseek flash 0731 or GPT 5.6 Luna, which already are almost free. Other than for things where the t/s are critical, it seems like a bad idea to etch a model into silicon.
That is fkin wild. o3 was just a year ago? The progress is truly insane.
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#173Earlier quoted context omitted.
Considering the rate of model development and rail hopping, seems like baking models into silicon is speed-running obsolescence.
I'd gladly pay for a Claude Opus 4.6 Thinking High in silicon and use it for 1-2 years. It's good enough for many coding tasks.
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#174Earlier quoted context omitted.
It's a terrible moat. You etch the silicon then nobody wants to run it in 6 months because models have advanced that much further.
OTOH, people get a new iPhone every year and they are ok with it.
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#175Earlier quoted context omitted.
A model can't be updated, and a chip that is only relevant for 6 months at max?
People already buy new phones every year, this just creates even more reason to do so
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#176The demo: https://chatjimmy.ai/
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#177Earlier quoted context omitted.
OTOH, people get a new iPhone every year and they are ok with it.
How is that in any way related to a consumer device? This method doesn't reduce physical memory requirements, so still results in huge die area. This isn't a for-end-user thing, probably for decades.
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#178Earlier quoted context omitted.
I'd gladly pay for a Claude Opus 4.6 Thinking High in silicon and use it for 1-2 years. It's good enough for many coding tasks.
It costs something like $300,000 for the hardware to run a model of that size. You'd pay that for a single model for 1-2 years? Not even the AI companies can justify that kind of spend which is why they keep extending the expected lifespan on their hardware in the accounting.
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#179Earlier quoted context omitted.
My question is what changes about LLM use cases when you’re getting 1000 tok/s? Models in silicon might dramatically change how we think about them.
That likely isn't as relevant for on-device iPhone usage as it is for Real Work™. I won't notice the difference between 50tps and 1000tps when asking Siri a question.
But if the basic premise of "good enough LLM at insane throughput" holds, I think it could qualitatively change local uses of LLMs. At a certain speed point, you're able to move from request -> response to a cascade of tool calling and "subagents", which could allow a small model to be much more useful, if provided with a lot of local data and tool calls.
That said, this is assuming you could stuff a "good enough" model into a phone with Taalas-like technology. The Taalas tech demo was an 8B parameter model and required hundreds of watts (IIRC) to run. The efficiency was good given the speed (as I understand), but it's not clear at all that the approach scales small enough to be a sensible coprocessor on an iPhone or whatever.
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#180Earlier quoted context omitted.
I freakin' love this demo. It feels magical.
I had the same reaction but then I showed it to my partner. She completely didn't get it, in her words "how can it be thinking of a good answer when it's that quick?" I tried to explain but I fear were probably going to be adding artificial sleeps to these things to convince the masses it's doing something clever.