Live data from Hacker News

AMD acquires Taalas to boost inference performance by etching models in silicon

theregister.com

271–280 of 725 posts

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#271

Thinking that five or six years from now, Fable-level intelligence could be provided at 100x the current speed... makes me feel lost. I cannot imagine what the future will look like.

It will be cool but also violent and terrible.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#273

Earlier quoted context omitted.

It costs something like $300,000 for the hardware to run a model of that size. You'd pay that for a single model for 1-2 years? Not even the AI companies can justify that kind of spend which is why they keep extending the expected lifespan on their hardware in the accounting.

> It costs something like $300,000 for the hardware to run a model of that size You did not compute that as the cost for a speculative card from Taalas, right?

It's the cost of the current nvidia hardware used to run these models. Of course all bets are off if you are accounting for some future chip that doesn't exist yet which could cost less.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#274

Earlier quoted context omitted.

My question is what changes about LLM use cases when you’re getting 1000 tok/s? Models in silicon might dramatically change how we think about them.

Did you use chatjimmy? It's somewhat terrifying to use when you think of the potential results with a better model. Ok, real life example: I now spend most of my time, as a developer, waiting for the agent to do its thing (after careful prompting, I'm also thinking about work stuff, don't worry I'm not useless). What if it gave back the same excellent results, but instantaneously? Why, then, I certainly would become…

I'm not sure inference speed is always the slowest thing for me right now. The agent is running tests, loading webpages, etc, which all take time. I don't know if a fast agent would speed things up in all cases.

That said, it obviously depends on the project.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#278

Thinking that five or six years from now, Fable-level intelligence could be provided at 100x the current speed... makes me feel lost. I cannot imagine what the future will look like.

Cerebras already runs large models like Kimi 2.6 or GLM at like 30x speed. 100 times is next year, not six years.

You can actually test it out on their website, just imagine 3 x faster and maybe 15% smarter.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#279

I'm surprised neither OpenAI nor Anthropic made this move first. The Chinese open weight models are pulling ahead and commoditizing their value proposition. Baking models onto silicon would've been the next logical move to get a moat. Google is already doing this and has an experimental project on top of already having TPUs and cramming their quantized flash onto individual TPUs for inference.

A model can't be updated, and a chip that is only relevant for 6 months at max?

Base model sure, but the stack will be hybrid. It’s still early days here. Too bad FPGAs have such large feature size.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#280

Earlier quoted context omitted.

"seems like baking models into silicon is speed-running obsolescence" Now maybe. When models are flying passenger aircraft, other prerogatives will assert themselves. When a 50TB ROM means you can impulse purchase a ChatGPT 6.3 xhigh that runs on batteries, yet more use cases will be apparent.

Or autonomous weapon systems, missiles, and drones.

Why would they need multi TB frontier models?
Post reply on HN