Live data from Hacker News

AMD acquires Taalas to boost inference performance by etching models in silicon

theregister.com

191–200 of 712 posts

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#191
post #174

Earlier quoted context omitted.

How is that in any way related to a consumer device? This method doesn't reduce physical memory requirements, so still results in huge die area. This isn't a for-end-user thing, probably for decades.

Ok, how long until nvidia gives us a new GPU?

I don't follow. How is that related? GPUs don't have fixed memory. You don't throw them away when you want to load a new model.

NVIDIA will probably give us a new GPU when someone competent in the free market decides they want wheelbarrows full of money. Unfortunately, AMD is entirely, incomprehensibly, incompetent, to the point where I can only assume they're colluding with Nvidia, behind the scenes.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#192
post #137

Field reprogrammable, it's an FPGA on steroids. Field upgradable. Burnt in, it needs a zif socket and easy access in every car, aircraft, a pull out slot in a phone, or it's new era planned obselescence.

Why not have some a device/hardware that programs itself on-boot.

Sort of a FPGA, that (electrically) arranges the connections on-boot, and then it's like a static inference chip.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#193

Is there any LLM from exactly one year ago that would be worth running? In Aug 2025 you had - OpenAI o3 - Opus 4.1 - Gemini 2.5 Pro - Grok 4 Even if those were almost free to run, you'd be way better off with Deepseek flash 0731 or GPT 5.6 Luna, which already are almost free. Other than for things where the t/s are critical, it seems like a bad idea to etch a model into silicon.

> Is there any LLM from exactly one year ago that would be worth running? Bad perspective: consider the correction: "when are thresholds of sought quality reached"? Hence: not "is there a 10yo from last year that could compete with the current 13yo", but "will there be a 30(?)yo from last year that could compete with the current 33(?)yo" ('(?)': the scale of yearly growth in the future is uncertain).

It's not just about it "being smart enough". It's about there being actual user demand when it needs to compete with the shiny new model.

A 10 year old iPhone is probably good enough, but is there demand for it? In a vacuum a 10 year old iPhone is good, but why would you pick it if you can have a current one for a reasonable price?

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#194

Earlier quoted context omitted.

My question is what changes about LLM use cases when you’re getting 1000 tok/s? Models in silicon might dramatically change how we think about them.

That likely isn't as relevant for on-device iPhone usage as it is for Real Work™. I won't notice the difference between 50tps and 1000tps when asking Siri a question.

But wouldn't higher tps allow for more reasoning or other hidden processes, potententially making a smarter model?

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#195

Earlier quoted context omitted.

My question is what changes about LLM use cases when you’re getting 1000 tok/s? Models in silicon might dramatically change how we think about them.

That likely isn't as relevant for on-device iPhone usage as it is for Real Work™. I won't notice the difference between 50tps and 1000tps when asking Siri a question.

That order of magnitude could be the difference between "the users wants me to open the notes app, let's open it" and "I've scanned all your notes before you could blink and found what you're looking for".

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#197

Earlier quoted context omitted.

I don't think this works out from a cost/silicon perspective. Small models already run pretty well in software (since the weights fit in cache) and big models require silicon area proportional to the size of weights. On a mobile device putting a chip like this is competing directly in BOM and power against a whole lot more l3 cache, and the l3 cache makes everything faster

My question is what changes about LLM use cases when you’re getting 1000 tok/s? Models in silicon might dramatically change how we think about them.

The best way I can explain it is that it's the same feeling when I upgraded from 56k dialup to cable broadband.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#198

Earlier quoted context omitted.

People already buy new phones every year, this just creates even more reason to do so

Outside of this website I've never met a person who buys a new phone every year. It's closer to every 3-4 years for most people.

Closer to every 5-6 years these days and with ram prices going up it will be even longer. Especially with the low/mid range phones, which are most phones outside some developed countries, people will keep their phones as long as they can.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#199

With web search and tool call a decent current generation model at the speed of the chatjimmy could do a lot. People saying it would be out of date are missing the point. It’s not going to make much sense for frontier companies that’s chasing the SOTA. But for a lot of business use cases if someone can put GLM 5.2 and sell it as a box, it would make so much sense. My partner has been asking for a “completely private”…

In my understanding the first Deep Think / Pro models were already very good as they were doing some kind of parallel repeated reasoning, thus were slow and expensive. So if chatjimmy speeds enables a fast deep think level performance, I think that would be great.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#200

I'm surprised neither OpenAI nor Anthropic made this move first. The Chinese open weight models are pulling ahead and commoditizing their value proposition. Baking models onto silicon would've been the next logical move to get a moat. Google is already doing this and has an experimental project on top of already having TPUs and cramming their quantized flash onto individual TPUs for inference.

You need to find customers for several-generations-ago models before this makes any sense. AMD is a lot more incentivized to look than mr vanilla llm is
Post reply on HN