Earlier quoted context omitted.
I believe this is a CPU/GPU vs ASIC comparison, rather than CPU vs GPU. They have always(ish) coexisted, being optimized for different things: ASICs have cost/speed/power advantages, but the design is more difficult than writing a computer program, and you can't reprogram them. Generally, you use an ASIC to perform a specific task. In this case, I think the takeaway is the LLM functionality here is performance-sensit…
It reminds me of the switch from GPUs to ASICs in bitcoin mining. I've been expecting this to happen.
How Taalas “prints” LLM onto a chip?
171–180 of 266 posts
Re: How Taalas “prints” LLM onto a chip?
#172Earlier quoted context omitted.
In a not-too-distant future (5 years?) small LLMs will be good enough to be used as generic models for most tasks. And if you have a dedicated ASIC small enough to fit in an iPhone, you have a truly local AI device with the bonus point that you get something really new to sell in every new generation (i.e. acces to an even more powerful model)
it doesn’t need to go in the phone if it only takes a few milliseconds to respond and is cheap
Re: How Taalas “prints” LLM onto a chip?
#173So if we assume this is the future, the useful life of many semiconductors will fall substantially. What part of the semiconductor supply chain would have pricing power in a world of producing many more different designs? Perhaps mask manufacturers?
It might be not that bad. “Good enough” open-weight models are almost there, the focus may shift to agentic workflows and effective prompting. The lifecycle of a model chip will be comparable to smartphones, getting longer and longer, with orchestration software being responsible for faster innovation cycles.
I distrust the notion. The bar of "good enough" seems to be bolted to "like today's frontier models", and frontier model performance only ever goes up.
Re: How Taalas “prints” LLM onto a chip?
#174Earlier quoted context omitted.
I'd be kind of shocked if Nvidia isn't playing with this. I don't expect it's like super commercially viable today, but for sure things need to trend to radically more efficient AI solutions.
These are chips that become e-waste the second a better a model comes out, and nvidia is already limited by TSMC capacity.
If you baked one of these into a smart speaker that could call tools to control lights and play music, it will still be able to do that when Llama 4 or 5 or 6 comes out.
Re: How Taalas “prints” LLM onto a chip?
#175Re: How Taalas “prints” LLM onto a chip?
#176I'm surprised people are surprised. Of course this is possible, and of course this is the future. This has been demonstrated already: why do you think we even have GPUs at all?! Because we did this exact same transition from running in software to largely running in hardware for all 2D and 3D Computer Graphics. And these LLMs are practically the same math, it's all just obvious and inevitable, if you're paying attent…
It's not certain this is the future: the obvious trade off is lack of flexibility: not only when a new model comes out, but also varying demand in the data centers - one day people want more LLM queries, another day more diffusion queries. Aaand, this blocks the holly grail of self improving models, beyond in-context learning. A realistic use case? More efficient vision based drone targeting in Ukraine/Taiwan/ whatev…
To your point, its neat tech, but the limitations are obvious since 'printing' only one LLM ensures further concentration of power. In other words, history repeats itself.
Re: How Taalas “prints” LLM onto a chip?
#177Re: How Taalas “prints” LLM onto a chip?
#178Earlier quoted context omitted.
I’m old enough to remember your typical computer filling warehouse-sized buildings. Nowadays, your average cellphone has more computing power than those behemoths. I have a micro SD card with 256GB capacity, and I think they are up to 2TB. On a device the size of a fingernail.
That is all definitely amazing, but data storage is a fundamentally different process with far fewer constraints than continuous computation.
Re: How Taalas “prints” LLM onto a chip?
#179Re: How Taalas “prints” LLM onto a chip?
#180I would appreciate some clarification on the "store 4 bits of data with one transistor" part. This doesn't sound remotely possible, but I am here to be convinced.
They declined to say: https://www.eetimes.com/taalas-specializes-to-extremes-for-e... Except they say it's fully digital, so not an analog multiplier