Live data from Hacker News

AMD acquires Taalas to boost inference performance by etching models in silicon

theregister.com

541–550 of 712 posts

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#541

The demo: https://chatjimmy.ai/

This is the only demo of 2026 which has blown my mind. If we can get to this speed with reasoning models, man... I can't even imagine the impact.

Looking at the history of technology, it's a question of when do we get there.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#542

I think what will happen is what happened to something like 4K video decoding before where it ends up in silicon costing almost nothing to run extremely fast on device. "Good enough" LLM functionality (for the use case) will be on-die or on-chip for cars, appliances, etc. This will provide speeds of chatjimmy at a battery-level power consumption. Probably this will also happen for software engineering. Some usb-power…

Scaling a model on a chip is quite hard. ChatJimmy is based on Llama 3.1 8 billion. Kimi K3 has 2.8 trillion parameters. That are 350x more parameters. I would expect that Gemma 4 E2B (approx 5.1 billion parameters) or maybe even Gemma 4 26 billion A4B at some point is running on a chip.

It's quite telling that the 8B Taalas chip was already reticle-sized on TSMC N6. I mean, we're talking about a process that does ~100 MTr/mm², ROM needs about one transistor per bit, but can probably be packed more densely than general logic. Something like, say, 150 megabit/mm² is not a lot. N6 has a 850 mm² reticle limit. This roughly tracks, the article says the chip has 8B parameters and apparently spends about half the area on ROM. There's a reason AI accelerators just use a ton of silicon area (each HBM3 die is >1000mm² of silicon). I imagine this is not terribly viable unless they make it a lot more space efficient e.g. using MLC ROM if they don't already, or use stacked dies with a ROM-optimized process. And then we're back to not cheap, though reticle chips were never in the cheap area to begin with.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#543

Earlier quoted context omitted.

How do humans?

Always the same trick of not answering the question and deflecting to „what about humans“. Can you folks not evaluate LLMs as the system they are, without vague gestures at how a different system behaves?

Evaluating LLMs is incredibly difficult. They are categorically different from any other system we have intuition about.

That said, I read the question I am replying to as a rhetorical one. If it was meant as a genuine question, curious about the question of meta knowledge, then I misread. Certainly the question is extremely interesting, for both LLMs and humans! But it's also obviously a very difficult one, as we don't even have a clear theory on how "knowing" works in the base case.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#544

Earlier quoted context omitted.

They are both the differentiator. AI previously provided speed but not quality. As soon as quality reached an acceptable threshold, the speed became the reigning factor. In my opinion the quality is still much lower, but speed means the cost is significantly lower also.

AI is already fast enough that human is a bottleneck. Hell, typing speed became a bottleneck like it was never before. I mean, if an agent can do half-decent work in less time than it takes the user to prompt them (and "user" in this context is a fast touch-typist like most programmers are), it's obvious it's not the agent that's the bottleneck anymore.

Because finishing someone else's (or something else's) "half decent work" to the point of "actually decent" becomes the bottleneck.

This has always been the case for human project management, and LLMs just aren't at that level yet.

It's more like everyone is speed running to how fast they can convince others that "half decent" is good enough. And for sure, newer models of LLM seem to be getting better at that.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#545

Earlier quoted context omitted.

Scaling a model on a chip is quite hard. ChatJimmy is based on Llama 3.1 8 billion. Kimi K3 has 2.8 trillion parameters. That are 350x more parameters. I would expect that Gemma 4 E2B (approx 5.1 billion parameters) or maybe even Gemma 4 26 billion A4B at some point is running on a chip.

It's quite telling that the 8B Taalas chip was already reticle-sized on TSMC N6. I mean, we're talking about a process that does ~100 MTr/mm², ROM needs about one transistor per bit, but can probably be packed more densely than general logic. Something like, say, 150 megabit/mm² is not a lot. N6 has a 850 mm² reticle limit. This roughly tracks, the article says the chip has 8B parameters and apparently spends about h…

Taalas exploits the low cardinality to store one 4-bit weight with one transistor. (They are using metal layer traces for the ROM, and connecting an access transistor to light up one of 16 options.)

Their system is honestly very efficient for the weights, the problem is the KV-cache. That's why HC1 only supports such short context, they use SRAM for that and spend most of what's left of the die for it. The recent advancements that made attention more efficient are probably going to be very useful for them.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#546
post #513

People are not talking enough how huge this is for robotics and IoT. Current robotics arhitectures are limited by tok/sec. How cares if its not a Fable model? This move undercuts NVIDIA directly.

With AI models like Mixture of Experts, many of those experts will be the real target here, as polished, refined and little to no change, they become fine candidates for being locked into silicon. Who knows, add some SRAM in there and small changes to those experts could be carried out without needing new silicon. Maybe AI models may become reduced to a collection of tiles you add to a chips one day, maybe sooner for…

An "Expert" is really just an unfortunate name for what amounts to a dense part of a sparse matrix and that's also an oversimplification.

It doesn't actually specialise in anything in particular that one can point to.

For this reason you can really transfer them between models.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#547

I think what will happen is what happened to something like 4K video decoding before where it ends up in silicon costing almost nothing to run extremely fast on device. "Good enough" LLM functionality (for the use case) will be on-die or on-chip for cars, appliances, etc. This will provide speeds of chatjimmy at a battery-level power consumption. Probably this will also happen for software engineering. Some usb-power…

Yes, it also opens up a faster recurring revenue model for hardware companies, faster model obsolescence than how often you change a computer, a server, or a GPU card. I hope they can figure out trillion-parameter models rapidly. Nvidia happened to be the best option for AI after building machines for graphics, so it makes sense they weren't the best idea from scratch for this specific use case. Especially given the scale of the demand and the possibility of recurring revenue, I hope a lot of smart people will try to solve it, compete with each other, and deliver us extremely fast and cheap intelligence.

And one can say LLMs are not as smart as a human, but a lot of the reasoning humans do for product and service generation isn't smart at all—it's just a bit of fuzzy input/output plus some reasoning rules. And then if you hook a robot up to the LLM, you can get results in atoms instead of bits.

I'm very excited about the future. I also hope it will stop money from flowing to bureaucrats who are incentived to keep the problems open to keep the money flowing, and instead facilitate sharing directly with the people (for example, no money to the state to solve homelessness—instead, spay instead with intelligence output to build a house and provide food as part of taxes.)

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#548

Earlier quoted context omitted.

It will be, I'm just not happy with my blog post before making it live. The blog will have a live interactive chat and a link to the repo with the HDL. I don't think anything I did was particularly novel, as I really just wanted to see how fast I could push a commodity FPGA to it's limit. Scaling to an ASIC or getting into the billions of params is where the real engineering is! This was just a side project for a sid…

What FPGA are you using? Why? What exactly did you implement? A full LLM? A subset of it, which collaborates with something running on CPU or GPU? Which LLM? Why? What language did you use to implement your thing: VHDL, Verilog, Vitis, something else? Why? I can think of at least 10 blog posts that I'd write before I write a single line of code. Publish early, publish soon ;-)

I'm using the AMD (Xilinx) K26. It's a Zynq Ultrascale+, the successor to the old classic 7000's. I'm running it on the KV260 dev board, because I'm using it for another side project.

The K26 has a quad core A53 core alongside the programmable logic (PL, or Fabric). The A53 is pretty weak, and doesn't have any hardware matmul operations, so despite the KV260 being sold as a 'vision ai starter kit' and the vitis object detection running on the arm cores, they're pretty weak cores for anything AI.

For my use, I need true determinism, so my vision pipeline is all implemented in the PL, and it was pretty disapointing that the vitis libraries are basically just opencv on linux, rather than really pushing the fabric. If I wanted probabalistic AI running on a CPU, then I sure as heck wouldn't choose a quad core A53.

Which led me to have a play with this, I saw the taalas/chatjimmy demo and wondered what I could push the fabric to.

The round trip time to DDR or CPU via AXI meant I had to keep the entire inference engine in fabric. The A53 is simply a pipe that gets a request from my server (which has a cloudflare tunnel to the real world for the live demo in the blog post) and manages a queue. So it feeds a string in, and gets a hopefully longer string back a few uS later.

It's all in verilog, because that's what i'm more used to. I did get Claude Code to do a moderate amount, as it's a side project on a side project after all, but pushing an FPGA to it's limit is definitely not as comfortable for it as it is writing a crud app in TS.

I'm using tinystories, as we are talking about megabytes of URAM/BRAM. If I used the DDR, it definitely would have been a real model, but that wasn't my goal. My goal was to hit 100,000tok/s, and even when I conceded on absolutely everything, with a token prediction size of 1tok, I topped out at 60,000tok/s. ?But increasing the window to make an actually plausible chat (story generator, it doesn't understand questions, you need to prompt it with 'once upon a time...' and it finishes it for example) I managed to break 20k tok/s.

I also created the lemmatised version, which was inspired by Kevin from the office (why use many word when few do trick) and trained a new model, I was expecting the output model to be smaller, but was suprised that it came out the same size, but it ran 30%ish faster. In hindsight it makes sense, the parameter count is fixed by the architecture, not the corpus, so training on compressed text doesn't shrink the model at all. What it does is compress the output distribution. The same story takes ~30% fewer characters to tell, so the effective speed goes up even though the per-token rate is identical. The dumbness is the optimisation.

Fully agree on publish early. The blog post with the live demo (a websocket straight to the board through a cloudflare tunnel, so you're genuinely talking to the fabric) is written and sitting in drafts while I fiddle with it. This thread is my peer pressure, it goes live in the next day or two.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#549

I think what will happen is what happened to something like 4K video decoding before where it ends up in silicon costing almost nothing to run extremely fast on device. "Good enough" LLM functionality (for the use case) will be on-die or on-chip for cars, appliances, etc. This will provide speeds of chatjimmy at a battery-level power consumption. Probably this will also happen for software engineering. Some usb-power…

I think you’re making another good point as well:

Specific traces for specific inferencing will mean that some generations get deprecated. Look at H265.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#550

Earlier quoted context omitted.

My question is what changes about LLM use cases when you’re getting 1000 tok/s? Models in silicon might dramatically change how we think about them.

It's pretty clear if you see what's happening on current phones. Autocorrect that works. Reply suggestions that almost work, just need to be tad more accurate (probably more of a data access issue than model) and a tad faster to look completely seamless. Screenshots with automated text detection and OCR and automatic interpretation (different suggested actions for when something on the picture looks like a web link,…

[dead]
Post reply on HN