8B coefficients are packed into 53B transistors, 6.5 transistors per coefficient. Two-inputs NAND gate takes 4 transistors and register takes about the same. One coefficient gets processed (multiplied by and result added to a sum) with less than two two-inputs NAND gates. I think they used block quantization: one can enumerate all possible blocks for all (sorted) permutations of coefficients and for each layer place…
How Taalas “prints” LLM onto a chip?
161–170 of 266 posts
Re: How Taalas “prints” LLM onto a chip?
#162Re: How Taalas “prints” LLM onto a chip?
#163Re: How Taalas “prints” LLM onto a chip?
#164Earlier quoted context omitted.
I'd be kind of shocked if Nvidia isn't playing with this. I don't expect it's like super commercially viable today, but for sure things need to trend to radically more efficient AI solutions.
These are chips that become e-waste the second a better a model comes out, and nvidia is already limited by TSMC capacity.
In the real world, theres talking refrigerators who dont need to know how to recite shakespeare.
Re: How Taalas “prints” LLM onto a chip?
#165[dead]
Re: How Taalas “prints” LLM onto a chip?
#166This would be a very interesting future. I can imagine Gemma 5 Mini running locally on hardware, or a hard-coded "AI core" like an ALU or media processor that supports particular encoding mechanisms like H.264, AV1, etc. Other than the obvious costs (but Taalas seems to be bringing back the structured ASIC era so costs shouldn't be that low [1]), I'm curious why this isn't getting much attention from larger companies…
Time is money and when you're competing with multiple companies with little margin for error you'll focus all your effort into releasing things quickly.
This chip is "only" a performance boost. It will unlock a lot of potential, but startups can't divide their attention like this. Big companies like google are surely already investigating this venue, but they might lack hardware expertise.
Re: How Taalas “prints” LLM onto a chip?
#167Earlier quoted context omitted.
I'd be kind of shocked if Nvidia isn't playing with this. I don't expect it's like super commercially viable today, but for sure things need to trend to radically more efficient AI solutions.
These are chips that become e-waste the second a better a model comes out, and nvidia is already limited by TSMC capacity.
Re: How Taalas “prints” LLM onto a chip?
#168Earlier quoted context omitted.
It's not certain this is the future: the obvious trade off is lack of flexibility: not only when a new model comes out, but also varying demand in the data centers - one day people want more LLM queries, another day more diffusion queries. Aaand, this blocks the holly grail of self improving models, beyond in-context learning. A realistic use case? More efficient vision based drone targeting in Ukraine/Taiwan/ whatev…
In a not-too-distant future (5 years?) small LLMs will be good enough to be used as generic models for most tasks. And if you have a dedicated ASIC small enough to fit in an iPhone, you have a truly local AI device with the bonus point that you get something really new to sell in every new generation (i.e. acces to an even more powerful model)
Re: How Taalas “prints” LLM onto a chip?
#169Earlier quoted context omitted.
These are chips that become e-waste the second a better a model comes out, and nvidia is already limited by TSMC capacity.
Only in VC backed funding land. In the real world, theres talking refrigerators who dont need to know how to recite shakespeare.
Re: How Taalas “prints” LLM onto a chip?
#170Does this mean computer boards will someday have one or more slots for an AI chip? Or peripheral devices containing AI models, which can be plugged into computer's high speed port?
Unless someone finds a way to turn these thijgs into a bios module.