Earlier quoted context omitted.
A cartridge slot for models is a fun idea. Instead of one chip running any model, you get one model or maybe a family of models per chip at (I assume) much better perf/watt. Curious whether the economics work out for consumer use or if this stays in the embedded/edge space.
Plug it into skull bone. Neuralink + slot for a model that you can buy in s grocery store instead of prepaid Netflix card.
How Taalas “prints” LLM onto a chip?
141–150 of 266 posts
Re: How Taalas “prints” LLM onto a chip?
#142Earlier quoted context omitted.
That slot is called USB-C. I can fully imagine inference ASICs coming in powerbank form factor that you'd just plug and play.
Like the chip-software in Gibson’s sprawl, from the micro-soft to the ROM cowboy to the Aleph, the endgame of computertool distribution is via single-use chunks of quasi-biological computronium
Re: How Taalas “prints” LLM onto a chip?
#143Earlier quoted context omitted.
The only product they've announced at the moment [0] is a PCI-e card. It's more like a small power bank than a big thumb drive. But sure, the next generation could be much smaller. It doesn't require battery cells, (much) heat management, or ruggedization, all of which put hard limits on how much you can miniaturise power banks. [0] https://taalas.com/the-path-to-ubiquitous-ai/
I’m old enough to remember your typical computer filling warehouse-sized buildings. Nowadays, your average cellphone has more computing power than those behemoths. I have a micro SD card with 256GB capacity, and I think they are up to 2TB. On a device the size of a fingernail.
Re: How Taalas “prints” LLM onto a chip?
#144I'm surprised people are surprised. Of course this is possible, and of course this is the future. This has been demonstrated already: why do you think we even have GPUs at all?! Because we did this exact same transition from running in software to largely running in hardware for all 2D and 3D Computer Graphics. And these LLMs are practically the same math, it's all just obvious and inevitable, if you're paying attent…
"This has been demonstrated already…" I think burning the weights into the gates is kinda new. ("Weights to gates." "Weighted gates"? "Gated weights"?)
Re: How Taalas “prints” LLM onto a chip?
#145[dead]
The network latency bit deserves more attention. I’ve been trying to find out where AI companies are physically serving LLMs from but it’s difficult to find information about this. If I’m sitting in London and use Claude, where are the requests actually being served? The ideal world would be an edge network like Cloudflare for LLMs so a nearby POP serves your requests. I’m not sure how viable this is. On classic hard…
Unfortunately, as with most of the AI providers, it's wherever they've been able to find available power and capacity. They've contracts with all of the large cloud vendors and lack of capacity is significant enough of an issue that locality isn't really part of the equation.
The only things they're particular about locality for is the infrastructure they use for training runs, where they need lots of interconnected capacity with low latency links.
Inference is wherever, whenever. You could be having your requests processed halfway around the world, or right next door, from one minute to the next.
Re: How Taalas “prints” LLM onto a chip?
#146I'm surprised people are surprised. Of course this is possible, and of course this is the future. This has been demonstrated already: why do you think we even have GPUs at all?! Because we did this exact same transition from running in software to largely running in hardware for all 2D and 3D Computer Graphics. And these LLMs are practically the same math, it's all just obvious and inevitable, if you're paying attent…
"This has been demonstrated already…" I think burning the weights into the gates is kinda new. ("Weights to gates." "Weighted gates"? "Gated weights"?)
Re: How Taalas “prints” LLM onto a chip?
#147Earlier quoted context omitted.
Not if you need 200w power to run inference.
USB-C can do up to 240W. These days I power all my devices with a USB hub, even my Lipo charger.
Re: How Taalas “prints” LLM onto a chip?
#148I'm surprised people are surprised. Of course this is possible, and of course this is the future. This has been demonstrated already: why do you think we even have GPUs at all?! Because we did this exact same transition from running in software to largely running in hardware for all 2D and 3D Computer Graphics. And these LLMs are practically the same math, it's all just obvious and inevitable, if you're paying attent…
Re: How Taalas “prints” LLM onto a chip?
#149I can imagine, where this becomes a mainstream PCIe extension card. Like back in days we had separate graphics card, audio card etc. Now AI card. So to upgrade the PC to latest model, we could buy a new card, load up the drivers and boom, intelligence upgrade of the PC. This would be so cool.
Another commenter mentioned how we keep cycling between local and server-based compute/storage as the dominant approach, and the cycle itself seems to be almost a law of nature. Nonetheless, regardless of where we're currently at in the cycle, there will always be both large and small players who want everything on-prem as much as possible.
Re: How Taalas “prints” LLM onto a chip?
#150Earlier quoted context omitted.
Apple should have done this yesterday. A local AI on my phone/Macbook is all I really want from this tech. The cloud-based AI (OpenAI, etc.) are todays AOL.
The die size is huge. This isn’t the kind of chip that would go into your MacBook, let alone an iPhone. It’s for cloud based servers.