Live data from Hacker News

How Taalas “prints” LLM onto a chip?

anuragk.com

231–240 of 266 posts

Re: How Taalas “prints” LLM onto a chip?

#231
post #3

This would be a very interesting future. I can imagine Gemma 5 Mini running locally on hardware, or a hard-coded "AI core" like an ALU or media processor that supports particular encoding mechanisms like H.264, AV1, etc. Other than the obvious costs (but Taalas seems to be bringing back the structured ASIC era so costs shouldn't be that low [1]), I'm curious why this isn't getting much attention from larger companies…

Apple should have done this yesterday. A local AI on my phone/Macbook is all I really want from this tech. The cloud-based AI (OpenAI, etc.) are todays AOL.

They did do it yesterday.

And it produced fake headlines and summaries including the threat of lawsuits from involved person(s).

Apple usually waits until somebody else has refined a technology to "invent" it, but I guess they couldn't wait for this one.

Re: How Taalas “prints” LLM onto a chip?

#232

Earlier quoted context omitted.

*the whole server uses 2.2kw or whatever, not a single board. I think that was for 8 boards or something.

Oh does it? Thanks for the clarification then. Their home page said 2.5kW so I assumed that's what it is. To be fair, 2.5kW does sound too much for a single 3x3cm chip, it would probably melt.

More powwwwaaa!

Yeah, though I suppose once we get properly 3d silicon I would not be surprised at power rating for that, 3cm^3 would be something to behold.

Re: How Taalas “prints” LLM onto a chip?

#234

Earlier quoted context omitted.

This is the same justification that was used to ship the (now almost entirely defunct) NPUs on Apple and Android devices alike. The A18 iPhone chip has 15b transistors for the GPU and CPU; the Taalas ASIC has 53b transistors dedicated to inference alone. If it's anything like NPUs, almost all vendors will bypass the baked-in silicon to use GPU acceleration past a certain point. It makes much more sense to ship a CUDA…

Why are you thinking about phones specifically? Most heavy users are on laptops and workstations. On smartphones there might be a few more innovations necessary (low latency AI computing on the edge?)

Many laptops and workstations also fell for the NPU meme, which in retrospect was a mistake compared to reworking your GPU architecture. Those NPUs are all dark silicon now, just like these Taalas chips will be in 12-24 months.

Dedicated inference ASICs are a dead end. You can't reprogram them, you can't finetune them, and they won't keep any of their resale value. Outside cruise missiles it's hard to imagine where such a disposable technology would be desirable.

Re: How Taalas “prints” LLM onto a chip?

#236
post #3

This would be a very interesting future. I can imagine Gemma 5 Mini running locally on hardware, or a hard-coded "AI core" like an ALU or media processor that supports particular encoding mechanisms like H.264, AV1, etc. Other than the obvious costs (but Taalas seems to be bringing back the structured ASIC era so costs shouldn't be that low [1]), I'm curious why this isn't getting much attention from larger companies…

Apple should have done this yesterday. A local AI on my phone/Macbook is all I really want from this tech. The cloud-based AI (OpenAI, etc.) are todays AOL.

https://developer.apple.com/documentation/FoundationModels

Re: How Taalas “prints” LLM onto a chip?

#238

Earlier quoted context omitted.

and run an outdated model for 3 years while progress is exponential? what is the point of that

When output is good enough, other considerations become more important. Most people on this planet cannot afford even an AI subscription, and cost of tokens is prohibitive to many low margin businesses. Privacy and personalization matter too, data sovereignty is a hot topic. Besides, we already see how focus has shifted to orchestration, which can be done on CPU and is cheap - software optimizations may compensate ha…

Taalas is more expensive than NPUs not less. You have GPU/NPU at home; just use it.

Re: How Taalas “prints” LLM onto a chip?

#239
post #3

This would be a very interesting future. I can imagine Gemma 5 Mini running locally on hardware, or a hard-coded "AI core" like an ALU or media processor that supports particular encoding mechanisms like H.264, AV1, etc. Other than the obvious costs (but Taalas seems to be bringing back the structured ASIC era so costs shouldn't be that low [1]), I'm curious why this isn't getting much attention from larger companies…

> I'm curious why this isn't getting much attention from larger companies. I can see two potential reasons: 1) Most of the big players seem convinced that AI is going to continue to improve at the rate it did in 2025, if their assumption is somehow correct by the time any chip entered mass production it would be obsolete. 2) The business model of the big players is to sell expensive subscriptions, and train on and se…

[deleted]

Re: How Taalas “prints” LLM onto a chip?

#240
post #103

> It took them two months, to develop chip for Llama 3.1 8B. In the AI world where one week is a year, it's super slow. But in a world of custom chips, this is supposed to be insanely fast. LLama 3.1 is like 2 years at this point. Taking two months to convert a model that only updates every 2 years is very fast

It only looks that way because Llama failed. Good models like Qwen are shipping every 6 months.
Post reply on HN