Live data from Hacker News

How Taalas “prints” LLM onto a chip?

anuragk.com

141–150 of 266 posts

Re: How Taalas “prints” LLM onto a chip?

#141
post #44

Earlier quoted context omitted.

A cartridge slot for models is a fun idea. Instead of one chip running any model, you get one model or maybe a family of models per chip at (I assume) much better perf/watt. Curious whether the economics work out for consumer use or if this stays in the embedded/edge space.

Plug it into skull bone. Neuralink + slot for a model that you can buy in s grocery store instead of prepaid Netflix card.

We better solve the energy usage and cooling first otherwise that will be a very spicy body mod.

Re: How Taalas “prints” LLM onto a chip?

#142
post #117

Earlier quoted context omitted.

That slot is called USB-C. I can fully imagine inference ASICs coming in powerbank form factor that you'd just plug and play.

Like the chip-software in Gibson’s sprawl, from the micro-soft to the ROM cowboy to the Aleph, the endgame of computertool distribution is via single-use chunks of quasi-biological computronium

Michael Bay just read "computronium" and spawned an 8 movie franchise in his head.

Re: How Taalas “prints” LLM onto a chip?

#143
post #31

Earlier quoted context omitted.

The only product they've announced at the moment [0] is a PCI-e card. It's more like a small power bank than a big thumb drive. But sure, the next generation could be much smaller. It doesn't require battery cells, (much) heat management, or ruggedization, all of which put hard limits on how much you can miniaturise power banks. [0] https://taalas.com/the-path-to-ubiquitous-ai/

I’m old enough to remember your typical computer filling warehouse-sized buildings. Nowadays, your average cellphone has more computing power than those behemoths. I have a micro SD card with 256GB capacity, and I think they are up to 2TB. On a device the size of a fingernail.

That is all definitely amazing, but data storage is a fundamentally different process with far fewer constraints than continuous computation.

Re: How Taalas “prints” LLM onto a chip?

#144

I'm surprised people are surprised. Of course this is possible, and of course this is the future. This has been demonstrated already: why do you think we even have GPUs at all?! Because we did this exact same transition from running in software to largely running in hardware for all 2D and 3D Computer Graphics. And these LLMs are practically the same math, it's all just obvious and inevitable, if you're paying attent…

"This has been demonstrated already…" I think burning the weights into the gates is kinda new. ("Weights to gates." "Weighted gates"? "Gated weights"?)

Geights? Wates?

Re: How Taalas “prints” LLM onto a chip?

#145
post #119

[dead]

The network latency bit deserves more attention. I’ve been trying to find out where AI companies are physically serving LLMs from but it’s difficult to find information about this. If I’m sitting in London and use Claude, where are the requests actually being served? The ideal world would be an edge network like Cloudflare for LLMs so a nearby POP serves your requests. I’m not sure how viable this is. On classic hard…

> The network latency bit deserves more attention. I’ve been trying to find out where AI companies are physically serving LLMs from but it’s difficult to find information about this. If I’m sitting in London and use Claude, where are the requests actually being served?

Unfortunately, as with most of the AI providers, it's wherever they've been able to find available power and capacity. They've contracts with all of the large cloud vendors and lack of capacity is significant enough of an issue that locality isn't really part of the equation.

The only things they're particular about locality for is the infrastructure they use for training runs, where they need lots of interconnected capacity with low latency links.

Inference is wherever, whenever. You could be having your requests processed halfway around the world, or right next door, from one minute to the next.

Re: How Taalas “prints” LLM onto a chip?

#146

I'm surprised people are surprised. Of course this is possible, and of course this is the future. This has been demonstrated already: why do you think we even have GPUs at all?! Because we did this exact same transition from running in software to largely running in hardware for all 2D and 3D Computer Graphics. And these LLMs are practically the same math, it's all just obvious and inevitable, if you're paying attent…

"This has been demonstrated already…" I think burning the weights into the gates is kinda new. ("Weights to gates." "Weighted gates"? "Gated weights"?)

Is this not effectively the same thing as a Bitcoin ASIC?

Re: How Taalas “prints” LLM onto a chip?

#147
post #76
post #58

Earlier quoted context omitted.

Not if you need 200w power to run inference.

USB-C can do up to 240W. These days I power all my devices with a USB hub, even my Lipo charger.

Have you seen a device that can supply 240w and act as a data host? Or is the 240w only from dedicated chargers?

Re: How Taalas “prints” LLM onto a chip?

#148

I'm surprised people are surprised. Of course this is possible, and of course this is the future. This has been demonstrated already: why do you think we even have GPUs at all?! Because we did this exact same transition from running in software to largely running in hardware for all 2D and 3D Computer Graphics. And these LLMs are practically the same math, it's all just obvious and inevitable, if you're paying attent…

It's not certain this is the future: the obvious trade off is lack of flexibility: not only when a new model comes out, but also varying demand in the data centers - one day people want more LLM queries, another day more diffusion queries. Aaand, this blocks the holly grail of self improving models, beyond in-context learning. A realistic use case? More efficient vision based drone targeting in Ukraine/Taiwan/ whatevers next. That's the place where energy efficiency, processing speed, and also weight is most critical. Not sure how heavy ASICS are though, bit they should be proportional to the model size. I heard many complaints about onboard AI 'not being there yet', and this may change it. Not listing middle east as there is no serious jamming problem there.

Re: How Taalas “prints” LLM onto a chip?

#149

I can imagine, where this becomes a mainstream PCIe extension card. Like back in days we had separate graphics card, audio card etc. Now AI card. So to upgrade the PC to latest model, we could buy a new card, load up the drivers and boom, intelligence upgrade of the PC. This would be so cool.

This is exactly what's going to happen. Assuming no civilization-crippling or Great Filter events, anyway. At this point I fail to see how it could go any other way. The path has already been traveled, and governments (along with many other large organizations) will demand this functionality for themselves, which will eventually have a consumer market as well.

Another commenter mentioned how we keep cycling between local and server-based compute/storage as the dominant approach, and the cycle itself seems to be almost a law of nature. Nonetheless, regardless of where we're currently at in the cycle, there will always be both large and small players who want everything on-prem as much as possible.

Re: How Taalas “prints” LLM onto a chip?

#150

Earlier quoted context omitted.

Apple should have done this yesterday. A local AI on my phone/Macbook is all I really want from this tech. The cloud-based AI (OpenAI, etc.) are todays AOL.

The die size is huge. This isn’t the kind of chip that would go into your MacBook, let alone an iPhone. It’s for cloud based servers.

And computers used to be the size of a room. I think they can get it to iPhone size in the future, this is an early prototype.
Post reply on HN