Live data from Hacker News

AMD acquires Taalas to boost inference performance by etching models in silicon

theregister.com

211–220 of 712 posts

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#211

Earlier quoted context omitted.

Outside of this website I've never met a person who buys a new phone every year. It's closer to every 3-4 years for most people.

Closer to every 5-6 years these days and with ram prices going up it will be even longer. Especially with the low/mid range phones, which are most phones outside some developed countries, people will keep their phones as long as they can.

Would depend on the income levels, but yeah, buying a new phone these days is entirely a non essential luxury. An iphone easily lasts 7 years so the moment money is tight, it's a very easy choice to not buy a new one.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#212

Earlier quoted context omitted.

I find speed alone would be a game changer for current models. I hardly find any task anymore that the current frontier models can't do with max reasoning after several rounds of feedback (provided sufficient instruction and the right harness). But waiting an hour or more for reasoning to finish is getting really cumbersome. If they could do the same in seconds (and for cheap of course), I'm pretty sure we'd pretty s…

Can you give some examples of these tasks that require an hour or more of reasoning?

The recent maths prompts did. The 'you should find a breakthrough' one was several blocks of reasoning, each taking 90 minutes or so

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#213
post #137

Field reprogrammable, it's an FPGA on steroids. Field upgradable. Burnt in, it needs a zif socket and easy access in every car, aircraft, a pull out slot in a phone, or it's new era planned obselescence.

It can just be pcie

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#214

I'm surprised neither OpenAI nor Anthropic made this move first. The Chinese open weight models are pulling ahead and commoditizing their value proposition. Baking models onto silicon would've been the next logical move to get a moat. Google is already doing this and has an experimental project on top of already having TPUs and cramming their quantized flash onto individual TPUs for inference.

Just to see how fast it is try chatjimmy.ai

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#215

Earlier quoted context omitted.

My question is what changes about LLM use cases when you’re getting 1000 tok/s? Models in silicon might dramatically change how we think about them.

That likely isn't as relevant for on-device iPhone usage as it is for Real Work™. I won't notice the difference between 50tps and 1000tps when asking Siri a question.

If Siri is using a 3T model in high reasoning mode to answer your question you will.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#216
post #214

I'm surprised neither OpenAI nor Anthropic made this move first. The Chinese open weight models are pulling ahead and commoditizing their value proposition. Baking models onto silicon would've been the next logical move to get a moat. Google is already doing this and has an experimental project on top of already having TPUs and cramming their quantized flash onto individual TPUs for inference.

Just to see how fast it is try chatjimmy.ai

Pretty incredible to see. It reminds me of when I first used the Groq chatbot, except in this case it's a full response instantly.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#217
post #183

Earlier quoted context omitted.

ASICs is what took over Bitcoin mining, cheaper in all ways, and lasts longer than Nvidia GPUs for inference.

> cheaper in all ways, Bitcoin mining doesn't have large memory requirements, but does have huge compute requirements. ASICs work great there because it's very straightforward to add some circuits for computing hashes. If you _also_ have to add many GB of memory, then suddenly ASICs will cost as much or more than comparable off-the-shelf hardware and they won't be faster unless you've also invested in huge memory ban…

My understanding is an ASIC can last 10+ years, where are Nvidia enterprise GPUs are rated for 5...

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#218

How's that jive with the fact that they're introducing a new model every other week?

The new model every week is not necessary at this point really. What if you could run opus 5 for the next couple years at 1/20 the cost?

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#219
post #205
post #201

I've had an endgame idea in mind for a while. Models, probably first open weight ones like Kimi K3 class, are etched into silicon like this and sold as cartridges almost like old school game cartridges. You buy a USB-C dongle that the cartridge goes into, or for data centers you have PCI cards that take these in slots.

Each cartridge costs $1,000. Do you still want it?

Me? Probably not. A business or a hoster, sure. There'd probably end up being an aftermarket in used cartridges with slightly older but still good models on them.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#220

How's that jive with the fact that they're introducing a new model every other week?

Pipeline the burn into silicon, lower the latency as much as you can, for the 10-100x operation cost it's worth it. Imagine if frontier models cost $5/mtok and the 2nd or 3rd tier models cost $5/billion tokens for 3-month-old models.
Post reply on HN