Thinking that five or six years from now, Fable-level intelligence could be provided at 100x the current speed... makes me feel lost. I cannot imagine what the future will look like.
AMD acquires Taalas to boost inference performance by etching models in silicon
271–280 of 731 posts
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#272Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#273Earlier quoted context omitted.
It costs something like $300,000 for the hardware to run a model of that size. You'd pay that for a single model for 1-2 years? Not even the AI companies can justify that kind of spend which is why they keep extending the expected lifespan on their hardware in the accounting.
> It costs something like $300,000 for the hardware to run a model of that size You did not compute that as the cost for a speculative card from Taalas, right?
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#274Earlier quoted context omitted.
My question is what changes about LLM use cases when you’re getting 1000 tok/s? Models in silicon might dramatically change how we think about them.
Did you use chatjimmy? It's somewhat terrifying to use when you think of the potential results with a better model. Ok, real life example: I now spend most of my time, as a developer, waiting for the agent to do its thing (after careful prompting, I'm also thinking about work stuff, don't worry I'm not useless). What if it gave back the same excellent results, but instantaneously? Why, then, I certainly would become…
That said, it obviously depends on the project.
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#275Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#276Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#277Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#278Thinking that five or six years from now, Fable-level intelligence could be provided at 100x the current speed... makes me feel lost. I cannot imagine what the future will look like.
You can actually test it out on their website, just imagine 3 x faster and maybe 15% smarter.
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#279I'm surprised neither OpenAI nor Anthropic made this move first. The Chinese open weight models are pulling ahead and commoditizing their value proposition. Baking models onto silicon would've been the next logical move to get a moat. Google is already doing this and has an experimental project on top of already having TPUs and cramming their quantized flash onto individual TPUs for inference.
A model can't be updated, and a chip that is only relevant for 6 months at max?
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#280Earlier quoted context omitted.
"seems like baking models into silicon is speed-running obsolescence" Now maybe. When models are flying passenger aircraft, other prerogatives will assert themselves. When a 50TB ROM means you can impulse purchase a ChatGPT 6.3 xhigh that runs on batteries, yet more use cases will be apparent.
Or autonomous weapon systems, missiles, and drones.