Live data from Hacker News

AMD acquires Taalas to boost inference performance by etching models in silicon

theregister.com

611–620 of 712 posts

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#611
post #388

Earlier quoted context omitted.

Did you use chatjimmy? It's somewhat terrifying to use when you think of the potential results with a better model. Ok, real life example: I now spend most of my time, as a developer, waiting for the agent to do its thing (after careful prompting, I'm also thinking about work stuff, don't worry I'm not useless). What if it gave back the same excellent results, but instantaneously? Why, then, I certainly would become…

Errors compound, and making 1000 wrong decisions per hour, will not result in something useful. Maybe you‘ve tried setting up guardrails for good design or architecture at some point? I think it’s simply not possible to do that. It would certainly be an accelerator for people who know exactly what they want. And it would remove multi tasking, which I‘d appreciate.

I don't have a great answer but you pose a great question.

Obviously a CTO is not going to walk away from the technology just because it's not good enough. That much more incentive for someone to create a powerful enough harness that can direct that power safely and productively. Like a nuclear core, we'll need to come up with the graphite rods and water tank. And if tokens are essentially free, why not, for every million tokens, spend 10x tokens on code review, testing, etc?

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#612
post #383

Earlier quoted context omitted.

But Claude Opus 4.6 is not really practical. Taalas' process seems targeted for edge models. Their proof of concept model, for example, is a heavily quantized version of Llama 3.1 8B and even then they acknowledge their custom 3-bit/6-bit representation causes model quality degradation. Taalas is going to have a tough time putting a trillion-parameter model on one conventional die. Their HC1 die is already near the m…

A really fast qwen-3.6-27B type of model could be useful. With a specialized harness and this speed I 'd expect it to find many applications. Implementing a coding plan is the minimum I can think of.

I'm sure life would find a way. I'd love to see what kind of power-harnesses people have to come up with to steer 16k tps QPU's (Qwen Processing Units) productively.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#613
post #121

Given the fast churn of the models, how does it work out? Won’t the silicon etched model already be 1 or more versions behind by the time the silicon comes out. Though if it’s cheap enough, there certainly can be a market for cheaper model inferences.

The thing is, right now it is exploding because we are at the beginning of it. At some point, it will plateau at a specific level, and not that much quality will be gained. There is however leaps to make for efficiency.

The same can be said about the CISC computer: yes, new processors introduce new instructions that do something slightly faster, you could still crunch that with an older processor. The real benefit comes in clock cycles (that's why Arm with a reduced set can compete with x86).

Also: there are myriads of models, for myriads of tasks. Not all have the same development gains as we see for general purpose AI. If you etch those, you reduce your bill by factors down.

It also democratises models: Instead of running them on a cloud server by some company, you can run them at home, for coding tasks, without the need of internet connection, etc.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#614

Is there any LLM from exactly one year ago that would be worth running? In Aug 2025 you had - OpenAI o3 - Opus 4.1 - Gemini 2.5 Pro - Grok 4 Even if those were almost free to run, you'd be way better off with Deepseek flash 0731 or GPT 5.6 Luna, which already are almost free. Other than for things where the t/s are critical, it seems like a bad idea to etch a model into silicon.

What if it was Fable 5 baked in?

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#615

I think what will happen is what happened to something like 4K video decoding before where it ends up in silicon costing almost nothing to run extremely fast on device. "Good enough" LLM functionality (for the use case) will be on-die or on-chip for cars, appliances, etc. This will provide speeds of chatjimmy at a battery-level power consumption. Probably this will also happen for software engineering. Some usb-power…

Now the fun part, how will having an LLM in my washing machine help anything

It can ponder the meaning of its existence.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#616
post #490

Earlier quoted context omitted.

I think the lifecycle for these chips could stretch far longer. If you're offering these models on a two year lifecycle, then you'd be able to stand up your top tier (wouldn't need to be frontier) at high speed. Run (for example) Kimi K3 on it and give it a brand name: AcmeAI Carbon Market it as your premier (only) model at high throughput. Two years later you stand up MSICs for the new state of the art with entirely…

I'm somewhat doubtful that we will be seeing something as large as Kimi K3 in silicon any time soon. This tech can definitely scale up from the current 8B prototype, but - at least as far as my limited understanding of the tech involved goes - you cannot just ASIC a trillion weights model due to physical size constraints. ___ Specification HC1 Model Llama 3.1 8B (hardwired) Process TSMC 6nm Die size 815mm² ___ So the…

This is an architectural limitation that may be overcome by how you bake the MoE (mixture-of-experts) onto silicon.

If you could manage a per-die expert somehow and keep the expert routing gate relatively fast (through an interposer interconnect or doing wafer-scale Cerebras type shit) you don't need to keep the whole thing on the same die. Small dies with one expert per die on an interposer, and a very tiny router might be sufficient.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#617
post #337

Earlier quoted context omitted.

Presumably wealth would concentrate upwards even if AI was never made.

Yes it's a function of the monetary system. Absurd amounts of debt only certain people can access.

So you're saying we need to abolish to monetary system?

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#618

I think what will happen is what happened to something like 4K video decoding before where it ends up in silicon costing almost nothing to run extremely fast on device. "Good enough" LLM functionality (for the use case) will be on-die or on-chip for cars, appliances, etc. This will provide speeds of chatjimmy at a battery-level power consumption. Probably this will also happen for software engineering. Some usb-power…

Now the fun part, how will having an LLM in my washing machine help anything

It can insult the plumbing with "you're a dumb pipe".

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#619

I think what will happen is what happened to something like 4K video decoding before where it ends up in silicon costing almost nothing to run extremely fast on device. "Good enough" LLM functionality (for the use case) will be on-die or on-chip for cars, appliances, etc. This will provide speeds of chatjimmy at a battery-level power consumption. Probably this will also happen for software engineering. Some usb-power…

Now the fun part, how will having an LLM in my washing machine help anything

same way that having wifi does. by providing no actionable value, but boosting marketing materials

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#620

Earlier quoted context omitted.

The Taalas chips are not physically small. And part of their secret (if you look at the design) is just locating a bunch of memory soldered on the edges ( I belive higher amounts of SRAM ? )

Baking the base models on to ROM makes a lot of economic sense. SRAM for the KV cache & fine-tunes, not so much. Sure you’d get incredible speeds but it’s not scalable from a die-size or cost perspective. Rather base model on ROM + KV cache on DRAM is much more scalable. Also this would work great for edge devices that have a 2-5 year lifecycle.

There's in-betweens; '1T SRAM' or eDRAM.

Of course, 1T SRAM isn't really SRAM, but my understanding is it doesn't require external refresh like eDRAM, is a bit easier to fab on-die than eDRAM, and is half the mm2 per Megabit compared to real SRAM (15% more die size than eDRAM)...

Post reply on HN