Live data from Hacker News

AMD acquires Taalas to boost inference performance by etching models in silicon

theregister.com

711–720 of 737 posts

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#711
post #590

Earlier quoted context omitted.

A full wafer like Cerebras is about 60x that, and N2P has about 3x the transistor density. So right now it's technically feasible to etch a 1.4 trillion parameter model. So roughly DeepSeek-V4-Pro class. Imagine that running a factory, for example.

Cerebras have special techniques to work around etching errors / bad cores on their wafers. This is possible since their wafers are effectively hundreds of identical copies of redundant cores. Can't do that for a globally unique model. Etching failure in that situation would be like brain-damage in a human, all sorts of weird effects would start appearing.

A few hundred bad bits/transistors in a trillion+ parameter model would compromise its abilities not one iota...the models are inherently lossy and resistant to "brain damage"...

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#712
post #561

Earlier quoted context omitted.

Scaling a model on a chip is quite hard. ChatJimmy is based on Llama 3.1 8 billion. Kimi K3 has 2.8 trillion parameters. That are 350x more parameters. I would expect that Gemma 4 E2B (approx 5.1 billion parameters) or maybe even Gemma 4 26 billion A4B at some point is running on a chip.

Scaling is definetly hard - but there is no absolute requirement to put huge flagship models into this technology (although it might be possible over time). A fairly dumb but FAST model has it's own totally distinct use-cases even if it can't be scaled in size. Think about a LLM-infused-Alexa where the response time is instant. Where you can request it looks at hotel options in Montreal, and it starts answering in ha…

And outside of idiotic demos, who exactly is going to ask an LLM to look at hotels in Montreal for them? This usecase has never made sense to me in the slightest

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#713

Earlier quoted context omitted.

> Where you can request it looks at hotel options in Montreal, and it starts answering in half a second Yet the answers will get outdated quickly whilst the silicon is fixed.

>Yet the answers will get outdated quickly whilst the silicon is fixed. Bro is living in 2020 before rag was widely introduced.

I was told updates require replacing at least two layers of metal, though not whole thing. Was that not accurate? Can you say more?

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#714

Earlier quoted context omitted.

Well, 50TB ROM Taalas HC1 style would be apparently a 400000b transistor system through a chip sized 2.5 meters on the side... :)

Yes, I know. This view is how these problems are always perceived, decade after decade, as our predecessors filled rooms with iron and silicon, unable to fathom that the equivalent capacity and power would be a portable device 20 years later. We're not at some end point in this process: the devices we have now will appear just a primitive in the years to come as a 10MB 5.25" Winchester drive appears to us now. One of…

> decade after decade

That can only work when there is physical capacity for improvement though.

> underappreciated effects of the AI boom and associated money is that it has strongly reinvigorated R&D in hardware

Yes, absolutely: but the point at this stage is more about finding new possibilities in hardware architecture than the improvement of what we had. So

> * That will yield what it has always yielded[:] orders of magnitude jumps in capacity and performance*

That will yield new and renewed hardware technologies.

(Already the distinction between SRAM and DRAM was overly specialistic before this boom - now it's on our mind as we know we need to "expand", "make cheap", "integrate" or find alternatives.)

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#715
post #561

Earlier quoted context omitted.

Scaling is definetly hard - but there is no absolute requirement to put huge flagship models into this technology (although it might be possible over time). A fairly dumb but FAST model has it's own totally distinct use-cases even if it can't be scaled in size. Think about a LLM-infused-Alexa where the response time is instant. Where you can request it looks at hotel options in Montreal, and it starts answering in ha…

And outside of idiotic demos, who exactly is going to ask an LLM to look at hotels in Montreal for them? This usecase has never made sense to me in the slightest

> who exactly is going to ask an LLM to look at hotels in Montreal for them

Anybody who has a specific informal query ("SELECT ... FROM ... WHERE has_carpark AND ... ORDER BY score(has_jacuzzi , walk_distance(...) ...) DESC") but does not want to research and cross the different scattered info himself (does not want to build the virtual DB himself).

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#716

Earlier quoted context omitted.

>Yet the answers will get outdated quickly whilst the silicon is fixed. Bro is living in 2020 before rag was widely introduced.

I was told updates require replacing at least two layers of metal, though not whole thing. Was that not accurate? Can you say more?

You are talking about two different things.

Yes, to update the blueprint for new models two layers will be updated. That is the NN.

To instead update the data on which to operate you could use a RAG to query.

(As in "the Pathfinder 2.0 NN is on the chip; the geodata is in the OpenGeoMaps dump-DB-nightly" - not really overlapping with LLM+RAG but may give an idea in a different scenario.)

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#717
post #561

Earlier quoted context omitted.

Scaling is definetly hard - but there is no absolute requirement to put huge flagship models into this technology (although it might be possible over time). A fairly dumb but FAST model has it's own totally distinct use-cases even if it can't be scaled in size. Think about a LLM-infused-Alexa where the response time is instant. Where you can request it looks at hotel options in Montreal, and it starts answering in ha…

And outside of idiotic demos, who exactly is going to ask an LLM to look at hotels in Montreal for them? This usecase has never made sense to me in the slightest

In the slightest? The user goes to their computer and types in "Montreal hotels" and Google comes up with no shortage of results, including a bit from their LLM. So that's already happening, but how do you narrow down the results from that initial search? Click around on Expedia for an hour? You probably know what you care about, just tell the LLM that you have dogs or are a vegan or whatever instead of wasting a bunch of time doing it by hand yourself.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#718

Earlier quoted context omitted.

> Where you can request it looks at hotel options in Montreal, and it starts answering in half a second Yet the answers will get outdated quickly whilst the silicon is fixed.

>Yet the answers will get outdated quickly whilst the silicon is fixed. Bro is living in 2020 before rag was widely introduced.

> before rag was widely introduced

And when did RAG start to work properly as a mature, reliable technology?

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#719
post #704

Earlier quoted context omitted.

Tell the washing machine. "Washlexa I'm putting my gym clothes in. It was a hard workout I sweated a lot!" and the washer now knows what to do. No more temp, length, 2nd rinse proxy controls.

I now picture people talking to their washing machine like the intro of american psycho.

I have seen smiling people talking to their connected car, replying with a dumb condescending bad-actor voice.

I had also seen Frankie Boyle in his parody of Knight Rider: "Michael, I am Kitt, your car, stop taking the medications, they want you to take them so you will be unable to talk to me"...

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#720
post #400

Earlier quoted context omitted.

Its a cool demo, but its gpt-3.5 level stupid, or worse. edit: Ok, I will self-apologize. Its apparently a 3B model. Mighty impressive for what it does.

How time flies. Just 3.5 years ago, gpt-3.5 was touted as almost AGI, we're all going to be replaced by machines and worst case they will kill us all. And here we are, not much later, and it serves as the benchmark for "stupid"..

> gpt-3.5 was touted as almost AGI

If you have the names, so we can record them in the History of Laughing Stocks...

Also remember maybe all those who shouted "that is really, precisely unintelligent".

Incidentally: also Sam Altman, in an interview with Lex Fridman, said "that is not proper intelligence but currently we do not know how to get there" - if very Altman refused to call it AGI...

Post reply on HN