Live data from Hacker News

AMD acquires Taalas to boost inference performance by etching models in silicon

theregister.com

361–370 of 712 posts

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#361
post #245
post #192

Earlier quoted context omitted.

Why not have some a device/hardware that programs itself on-boot. Sort of a FPGA, that (electrically) arranges the connections on-boot, and then it's like a static inference chip.

FPGAs already configure themselves on boot.

I asked a LLM after posting my comment, to see if I had a genius idea or not,just for it to tell me the same as you, that's now they work already...

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#362

I've been eagerly awaiting their 2nd gen HC2, which uses multiple chips to host a "mid sized reasoning" [1] model. Its due in summer according to the article, I wonder if it will ever be released in that form now. [1] https://www.forbes.com/sites/karlfreund/2026/02/19/taalas-la...

> I wonder if it will ever be released in that form now.

Yeah, I had the same thought. The key thing for them was the price point at which they could deliver a ~30B model. I would buy one today if it was ~1000$ and could run whatever the best 30B model is today, at those speeds advertised. Even if the model becomes superseded by model.5 in a few months, there's still a lot of things you can do with a "good enough" model for some tasks. And things like maj@x or generate 10 times and choose "at a glance" what you like (think frontend stuff) would be worth it.

No idea if them selling to AMD is good or bad.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#363

Earlier quoted context omitted.

My question is what changes about LLM use cases when you’re getting 1000 tok/s? Models in silicon might dramatically change how we think about them.

Did you use chatjimmy? It's somewhat terrifying to use when you think of the potential results with a better model. Ok, real life example: I now spend most of my time, as a developer, waiting for the agent to do its thing (after careful prompting, I'm also thinking about work stuff, don't worry I'm not useless). What if it gave back the same excellent results, but instantaneously? Why, then, I certainly would become…

This is the same pitch that people make about AI today. Speed isn’t the differentiator, quality is

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#364
post #330

Earlier quoted context omitted.

I'd gladly pay for a Claude Opus 4.6 Thinking High in silicon and use it for 1-2 years. It's good enough for many coding tasks.

Not so long ago, I was good enough for many coding tasks. But I found that things can change in a hurry. Yes, a cheap and fast Opus4.6 can drive a lot of value in current context. But if we continue to craft bigger-and-bigger balls of mud, Opus 4.6 may end up hitting its conceptual ceiling and unable to contribute. Winding the clock back on your statement gives: > I'd gladly pay for a Claude Sonnet 3.5 in silicon and…

Assuming moore's law like progress, which I'm 100% sure isn't going to happen - I think we're at the top of the S curve already. But assuming dramatically increased intelligence every year this is still the exact same position as anyone who bought a computer in the last 5 decades. Yet, people did very much buy computers.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#366

Thinking that five or six years from now, Fable-level intelligence could be provided at 100x the current speed... makes me feel lost. I cannot imagine what the future will look like.

It tells me that they have some kind of insider knowledge that the models have hit their limits and won't be getting much better, and it makes sense economically speaking to just bake the current models and use them for the next 5-10 years. Looks like we're near the top of the S curve.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#367
post #278

Earlier quoted context omitted.

Cerebras already runs large models like Kimi 2.6 or GLM at like 30x speed. 100 times is next year, not six years. You can actually test it out on their website, just imagine 3 x faster and maybe 15% smarter.

Cerebras is literally the entire wafer, so it can't get bigger. So where is the jump from 30x to 100x coming from? Node improvements only yield like 10-20% gains these days...

They have a next generation, I don't really know if it will be 3 x or what but I heard it was significantly better.

Also there are other people innovating in hardware.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#369
post #366

Thinking that five or six years from now, Fable-level intelligence could be provided at 100x the current speed... makes me feel lost. I cannot imagine what the future will look like.

It tells me that they have some kind of insider knowledge that the models have hit their limits and won't be getting much better, and it makes sense economically speaking to just bake the current models and use them for the next 5-10 years. Looks like we're near the top of the S curve.

What would possibly tell you that?

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#370

Earlier quoted context omitted.

Cerebras chips are massive and do have more on the edge but they dont have any top or bottom cache do they?

They can't due to power density, I believe - they have to be run in a sandwiched waterblock with massive cooling, as far as I can tell. That's the biggest thing that baked weights gets you - a relatively modest watts-per-square-mm compare to cerebras, where they had to engineer a whole system to get the watts out of the chip

Do you think there's room for reducing power requirements? Obviously shrinking the process is a win, but is the existing implementation a "just make it work" phase that has opportunities to increase computational efficiency?

With that kind of speed and if even lower power requirements, they could release mini compute units with USB4/Thunderbolt for plug and play inference.

Post reply on HN