Earlier quoted context omitted.
Why not have some a device/hardware that programs itself on-boot. Sort of a FPGA, that (electrically) arranges the connections on-boot, and then it's like a static inference chip.
FPGAs already configure themselves on boot.
AMD acquires Taalas to boost inference performance by etching models in silicon
361–370 of 712 posts
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#362I've been eagerly awaiting their 2nd gen HC2, which uses multiple chips to host a "mid sized reasoning" [1] model. Its due in summer according to the article, I wonder if it will ever be released in that form now. [1] https://www.forbes.com/sites/karlfreund/2026/02/19/taalas-la...
Yeah, I had the same thought. The key thing for them was the price point at which they could deliver a ~30B model. I would buy one today if it was ~1000$ and could run whatever the best 30B model is today, at those speeds advertised. Even if the model becomes superseded by model.5 in a few months, there's still a lot of things you can do with a "good enough" model for some tasks. And things like maj@x or generate 10 times and choose "at a glance" what you like (think frontend stuff) would be worth it.
No idea if them selling to AMD is good or bad.
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#363Earlier quoted context omitted.
My question is what changes about LLM use cases when you’re getting 1000 tok/s? Models in silicon might dramatically change how we think about them.
Did you use chatjimmy? It's somewhat terrifying to use when you think of the potential results with a better model. Ok, real life example: I now spend most of my time, as a developer, waiting for the agent to do its thing (after careful prompting, I'm also thinking about work stuff, don't worry I'm not useless). What if it gave back the same excellent results, but instantaneously? Why, then, I certainly would become…
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#364Earlier quoted context omitted.
I'd gladly pay for a Claude Opus 4.6 Thinking High in silicon and use it for 1-2 years. It's good enough for many coding tasks.
Not so long ago, I was good enough for many coding tasks. But I found that things can change in a hurry. Yes, a cheap and fast Opus4.6 can drive a lot of value in current context. But if we continue to craft bigger-and-bigger balls of mud, Opus 4.6 may end up hitting its conceptual ceiling and unable to contribute. Winding the clock back on your statement gives: > I'd gladly pay for a Claude Sonnet 3.5 in silicon and…
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#365Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#366Thinking that five or six years from now, Fable-level intelligence could be provided at 100x the current speed... makes me feel lost. I cannot imagine what the future will look like.
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#367Earlier quoted context omitted.
Cerebras already runs large models like Kimi 2.6 or GLM at like 30x speed. 100 times is next year, not six years. You can actually test it out on their website, just imagine 3 x faster and maybe 15% smarter.
Cerebras is literally the entire wafer, so it can't get bigger. So where is the jump from 30x to 100x coming from? Node improvements only yield like 10-20% gains these days...
Also there are other people innovating in hardware.
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#368Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#369Thinking that five or six years from now, Fable-level intelligence could be provided at 100x the current speed... makes me feel lost. I cannot imagine what the future will look like.
It tells me that they have some kind of insider knowledge that the models have hit their limits and won't be getting much better, and it makes sense economically speaking to just bake the current models and use them for the next 5-10 years. Looks like we're near the top of the S curve.
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#370Earlier quoted context omitted.
Cerebras chips are massive and do have more on the edge but they dont have any top or bottom cache do they?
They can't due to power density, I believe - they have to be run in a sandwiched waterblock with massive cooling, as far as I can tell. That's the biggest thing that baked weights gets you - a relatively modest watts-per-square-mm compare to cerebras, where they had to engineer a whole system to get the watts out of the chip
With that kind of speed and if even lower power requirements, they could release mini compute units with USB4/Thunderbolt for plug and play inference.