Live data from Hacker News

AMD acquires Taalas to boost inference performance by etching models in silicon

theregister.com

681–690 of 710 posts

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#681
post #536

Earlier quoted context omitted.

Closer to every 5-6 years these days and with ram prices going up it will be even longer. Especially with the low/mid range phones, which are most phones outside some developed countries, people will keep their phones as long as they can.

I mean are there even any reasons to buy a new phone? If I compare the Pixel 6 Pro I'm using at the moment to current models, they are functionally identical. The only reason to upgrade might be getting a fresh battery and access to firmware updates. Otherwise I'd be happy to continue using it for the next 10 years.

I only buy phones if the current one starts showing signs of deterioration, mostly battery.

I swear a midrange Chinese phone from 2017 would be enough for me in 2026 to read HN/Whatsapp and some Youtube.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#682
post #410

Earlier quoted context omitted.

What would possibly tell you that?

Expensive as fuck to make chips, only makes sense if you believe whatever model you're creating a chip out of will not become completely irrelevant in 5-10 years.

That’s not what we are empirically seeing though. There is also no way of knowing that things can’t be progressing for the next 5 years. No foundational model company would be spending the amount they are on research if they thought it wasn’t going to pan out

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#683

Earlier quoted context omitted.

> I now spend most of my time, as a developer, waiting for the agent to do its thing (after careful prompting, I'm also thinking about work stuff, don't worry I'm not useless). you need to launch 10-15 more terminals, who is waiting these days? :)

You sound like my boss! I'm not really into the whole "burnout" thing though.

how can you get burned out just watching the work being done for you?? :)

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#684

So I guess, 1. This chip for an 8B model even if it was done at 5nm would still be twice the size of a conventional CPU die so what are the yields for this going to be like for even a 30B model? 2. They say 2 months but llama 3.1 was released 2024, ~2 years which is normal lead time for silicon, I suspect this would take longer if the architecture is not llama? 3. Can google do the same thing in house with their Gemm…

Exactly, what will be the size of big models? Maybe they aren’t targeting big models but where do they give size estimates?

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#685
post #235

What does it take to go from here to a model on a pcie card or an m.2 card, so I can plug one into my workstation / laptop? Will 'intelligence' become much like a gpu, where most people just live with the performance of whatever they have installed, outside large companies that must have cutting edge, or prosumers that have a incrementally better version than the masses? Are we a couple years away, a decade away, or…

> What does it take to go from here to a model on a pcie card or an m.2 card It is already that. > Will "intelligence" become much like a gpu As an option among the implementations. > Are we a couple years away They could mass produce now, but it makes no sense at this rate of improvements in the models.

Thanks for answering. This is an 8b model, which are mainly curiosities outside niche tasks. I guess I am asking how far we are away from having today's more generally useful frontier model equivalents widely available for everyday users in their personal pcs/laptops via a single pcie or m.2 drop in.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#686

I think what will happen is what happened to something like 4K video decoding before where it ends up in silicon costing almost nothing to run extremely fast on device. "Good enough" LLM functionality (for the use case) will be on-die or on-chip for cars, appliances, etc. This will provide speeds of chatjimmy at a battery-level power consumption. Probably this will also happen for software engineering. Some usb-power…

Now the fun part, how will having an LLM in my washing machine help anything

Never mind washing machine, how about a missile or a drone?

Or a bomb.. "let there be light"

https://www.youtube.com/watch?v=h73PsFKtIck

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#687

I think what will happen is what happened to something like 4K video decoding before where it ends up in silicon costing almost nothing to run extremely fast on device. "Good enough" LLM functionality (for the use case) will be on-die or on-chip for cars, appliances, etc. This will provide speeds of chatjimmy at a battery-level power consumption. Probably this will also happen for software engineering. Some usb-power…

Scaling a model on a chip is quite hard. ChatJimmy is based on Llama 3.1 8 billion. Kimi K3 has 2.8 trillion parameters. That are 350x more parameters. I would expect that Gemma 4 E2B (approx 5.1 billion parameters) or maybe even Gemma 4 26 billion A4B at some point is running on a chip.

Consider a model like https://huggingface.co/nvidia/NVIDIA-Nemotron-Parse-2.0 which just came out. 0.9B parameters and very accurate for doing a very specific task: document features classification.

Now imagine you have a chip which is just that model, but can do it at absolutely insane speed. Like tens of thousands of documents a second.

Same for things like text-to-speech or speech-to-text. Think of the accessibility wins if subtitling becomes insanely accurate and fast and omnipresent.

There are all sorts of domains like that, and the trend has been such that smaller models are getting smarter and smarter. If you can stick them in parking meters, traffic lights / street crossings, mobility aids, etc etc I just see so much potential win.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#688

Earlier quoted context omitted.

Slightly besides your point, but it's interesting how many here naturally ponder about how the current winner could or "should" keep winning, instead of how another company could become a competitor by doing the more clever thing the incumbent isn't thinking about.

It is not a “should”. At least not in the “we wish it were so” sense. It is more that there are multiple reasons why this idea (burning an LLM into silicone and deploying it into a device in people’s pockets) requires huge piles of cash and the kind of engineering chops only a few company posesses. Of course i would like it if a small upstart would do this, but it doesn’t seem likely as a posibility. They won’t have…

Good point. IMHO we're thinking of too much in single entities doing everything.

Right now it's kinda the world we live in, with Apple or Google doing the total vertical integration from chip design to retail shops, but it doesn't have to be that way.

It could be done the traditional way with for instance a joint venture receiving funds and expertise from several players in each of the field and collaborating with external companies to get to the final package.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#689

I feel like this will be the end of Taalas, AMD has for the most part of its history always chosen the wrong options.

always? like amd64 vs itanium?

(and AMD is still serious challenging Intel in x86 and GPUs today)

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#690
post #590

Earlier quoted context omitted.

It's quite telling that the 8B Taalas chip was already reticle-sized on TSMC N6. I mean, we're talking about a process that does ~100 MTr/mm², ROM needs about one transistor per bit, but can probably be packed more densely than general logic. Something like, say, 150 megabit/mm² is not a lot. N6 has a 850 mm² reticle limit. This roughly tracks, the article says the chip has 8B parameters and apparently spends about h…

A full wafer like Cerebras is about 60x that, and N2P has about 3x the transistor density. So right now it's technically feasible to etch a 1.4 trillion parameter model. So roughly DeepSeek-V4-Pro class. Imagine that running a factory, for example.

Cerebras have special techniques to work around etching errors / bad cores on their wafers. This is possible since their wafers are effectively hundreds of identical copies of redundant cores. Can't do that for a globally unique model.

Etching failure in that situation would be like brain-damage in a human, all sorts of weird effects would start appearing.

Post reply on HN