Earlier quoted context omitted.
A full wafer like Cerebras is about 60x that, and N2P has about 3x the transistor density. So right now it's technically feasible to etch a 1.4 trillion parameter model. So roughly DeepSeek-V4-Pro class. Imagine that running a factory, for example.
Cerebras have special techniques to work around etching errors / bad cores on their wafers. This is possible since their wafers are effectively hundreds of identical copies of redundant cores. Can't do that for a globally unique model. Etching failure in that situation would be like brain-damage in a human, all sorts of weird effects would start appearing.
AMD acquires Taalas to boost inference performance by etching models in silicon
711–720 of 737 posts
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#712Earlier quoted context omitted.
Scaling a model on a chip is quite hard. ChatJimmy is based on Llama 3.1 8 billion. Kimi K3 has 2.8 trillion parameters. That are 350x more parameters. I would expect that Gemma 4 E2B (approx 5.1 billion parameters) or maybe even Gemma 4 26 billion A4B at some point is running on a chip.
Scaling is definetly hard - but there is no absolute requirement to put huge flagship models into this technology (although it might be possible over time). A fairly dumb but FAST model has it's own totally distinct use-cases even if it can't be scaled in size. Think about a LLM-infused-Alexa where the response time is instant. Where you can request it looks at hotel options in Montreal, and it starts answering in ha…
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#713Earlier quoted context omitted.
> Where you can request it looks at hotel options in Montreal, and it starts answering in half a second Yet the answers will get outdated quickly whilst the silicon is fixed.
>Yet the answers will get outdated quickly whilst the silicon is fixed. Bro is living in 2020 before rag was widely introduced.
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#714Earlier quoted context omitted.
Well, 50TB ROM Taalas HC1 style would be apparently a 400000b transistor system through a chip sized 2.5 meters on the side... :)
Yes, I know. This view is how these problems are always perceived, decade after decade, as our predecessors filled rooms with iron and silicon, unable to fathom that the equivalent capacity and power would be a portable device 20 years later. We're not at some end point in this process: the devices we have now will appear just a primitive in the years to come as a 10MB 5.25" Winchester drive appears to us now. One of…
That can only work when there is physical capacity for improvement though.
> underappreciated effects of the AI boom and associated money is that it has strongly reinvigorated R&D in hardware
Yes, absolutely: but the point at this stage is more about finding new possibilities in hardware architecture than the improvement of what we had. So
> * That will yield what it has always yielded[:] orders of magnitude jumps in capacity and performance*
That will yield new and renewed hardware technologies.
(Already the distinction between SRAM and DRAM was overly specialistic before this boom - now it's on our mind as we know we need to "expand", "make cheap", "integrate" or find alternatives.)
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#715Earlier quoted context omitted.
Scaling is definetly hard - but there is no absolute requirement to put huge flagship models into this technology (although it might be possible over time). A fairly dumb but FAST model has it's own totally distinct use-cases even if it can't be scaled in size. Think about a LLM-infused-Alexa where the response time is instant. Where you can request it looks at hotel options in Montreal, and it starts answering in ha…
And outside of idiotic demos, who exactly is going to ask an LLM to look at hotels in Montreal for them? This usecase has never made sense to me in the slightest
Anybody who has a specific informal query ("SELECT ... FROM ... WHERE has_carpark AND ... ORDER BY score(has_jacuzzi , walk_distance(...) ...) DESC") but does not want to research and cross the different scattered info himself (does not want to build the virtual DB himself).
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#716Earlier quoted context omitted.
>Yet the answers will get outdated quickly whilst the silicon is fixed. Bro is living in 2020 before rag was widely introduced.
I was told updates require replacing at least two layers of metal, though not whole thing. Was that not accurate? Can you say more?
Yes, to update the blueprint for new models two layers will be updated. That is the NN.
To instead update the data on which to operate you could use a RAG to query.
(As in "the Pathfinder 2.0 NN is on the chip; the geodata is in the OpenGeoMaps dump-DB-nightly" - not really overlapping with LLM+RAG but may give an idea in a different scenario.)
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#717Earlier quoted context omitted.
Scaling is definetly hard - but there is no absolute requirement to put huge flagship models into this technology (although it might be possible over time). A fairly dumb but FAST model has it's own totally distinct use-cases even if it can't be scaled in size. Think about a LLM-infused-Alexa where the response time is instant. Where you can request it looks at hotel options in Montreal, and it starts answering in ha…
And outside of idiotic demos, who exactly is going to ask an LLM to look at hotels in Montreal for them? This usecase has never made sense to me in the slightest
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#718Earlier quoted context omitted.
> Where you can request it looks at hotel options in Montreal, and it starts answering in half a second Yet the answers will get outdated quickly whilst the silicon is fixed.
>Yet the answers will get outdated quickly whilst the silicon is fixed. Bro is living in 2020 before rag was widely introduced.
And when did RAG start to work properly as a mature, reliable technology?
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#719Earlier quoted context omitted.
Tell the washing machine. "Washlexa I'm putting my gym clothes in. It was a hard workout I sweated a lot!" and the washer now knows what to do. No more temp, length, 2nd rinse proxy controls.
I now picture people talking to their washing machine like the intro of american psycho.
I had also seen Frankie Boyle in his parody of Knight Rider: "Michael, I am Kitt, your car, stop taking the medications, they want you to take them so you will be unable to talk to me"...
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#720Earlier quoted context omitted.
Its a cool demo, but its gpt-3.5 level stupid, or worse. edit: Ok, I will self-apologize. Its apparently a 3B model. Mighty impressive for what it does.
How time flies. Just 3.5 years ago, gpt-3.5 was touted as almost AGI, we're all going to be replaced by machines and worst case they will kill us all. And here we are, not much later, and it serves as the benchmark for "stupid"..
If you have the names, so we can record them in the History of Laughing Stocks...
Also remember maybe all those who shouted "that is really, precisely unintelligent".
Incidentally: also Sam Altman, in an interview with Lex Fridman, said "that is not proper intelligence but currently we do not know how to get there" - if very Altman refused to call it AGI...