Now that I think about it, real time AI video might be a clear case of “You scientists were so preoccupied with whether you could or not, you forgot to ask if you should”
AMD acquires Taalas to boost inference performance by etching models in silicon
641–650 of 710 posts
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#642Earlier quoted context omitted.
"seems like baking models into silicon is speed-running obsolescence" Now maybe. When models are flying passenger aircraft, other prerogatives will assert themselves. When a 50TB ROM means you can impulse purchase a ChatGPT 6.3 xhigh that runs on batteries, yet more use cases will be apparent.
Well, 50TB ROM Taalas HC1 style would be apparently a 400000b transistor system through a chip sized 2.5 meters on the side... :)
One of the underappreciated effects of the AI boom and associated money is that it has strongly reinvigorated R&D in hardware: it is clear that there is a real application for far greater density and lower power demand, and people are now pursuing this much harder than they had been. That will yield what it has always yielded; orders of magnitude jumps in capacity and performance.
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#643I see a lot of discourse about it being fast-to-deprecation. But I see it a different way personally. Modern LLMs are trying to do more with less. Focus on doing the right thing the first time. Even if we squeeze dumb LLMs, the significantly faster speed means quicker iterations. So a bad decision doesn't cost the time and inference costs that it cost before. It theoretically changes the scale of errant token spend.…
And also, there are lots of tasks where models today are fine with doing. If you think of these things like appliances, who cares if it's not quite as powerful as the next generation? It was purchased to do a task, it still does that task very well. It feels like being in the 90s and asking "why buy a server today when they're going to be faster next year? Just keep renting mainframe time." Well maybe I just need a box to run our HR and payroll system, and this box manages to run it fine today.
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#644Earlier quoted context omitted.
Now the fun part, how will having an LLM in my washing machine help anything
You load it, tell it what's in and it sets the program, tells you what it set and why. You approve and off it goes
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#645Given the fast churn of the models, how does it work out? Won’t the silicon etched model already be 1 or more versions behind by the time the silicon comes out. Though if it’s cheap enough, there certainly can be a market for cheaper model inferences.
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#646Earlier quoted context omitted.
Taalas exploits the low cardinality to store one 4-bit weight with one transistor. (They are using metal layer traces for the ROM, and connecting an access transistor to light up one of 16 options.) Their system is honestly very efficient for the weights, the problem is the KV-cache. That's why HC1 only supports such short context, they use SRAM for that and spend most of what's left of the die for it. The recent adv…
Do you really need a KV cache if inference is that fast though?
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#647Earlier quoted context omitted.
Slightly besides your point, but it's interesting how many here naturally ponder about how the current winner could or "should" keep winning, instead of how another company could become a competitor by doing the more clever thing the incumbent isn't thinking about.
It is not a “should”. At least not in the “we wish it were so” sense. It is more that there are multiple reasons why this idea (burning an LLM into silicone and deploying it into a device in people’s pockets) requires huge piles of cash and the kind of engineering chops only a few company posesses. Of course i would like it if a small upstart would do this, but it doesn’t seem likely as a posibility. They won’t have…
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#648I'm surprised there's not more discussion about potential inflection points here. When technology gets faster, it opens up whole new classes of UX that were hard to predict For example, faster internet didn't mean being able to view 100x as many HTML4 web pages. It brought SaaS, streaming media and interactivity. I'm not good at predicting, but some ideas: 1. All information gets augmented in real time with personali…
How about an actually smart thermostat that checks the weather and possibly makes decisions more like a human would i.e. tool calling, judgement, preference, history, personal plans.
Sure we have thermostats and you can configure rules and data sources, hook up Google calendar, etc but it has to be all predetermined and breaks as soon as anything stops working. AI could make this less brittle. AI agents are more flexible.
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#649Earlier quoted context omitted.
4. Pervasive , distributed dragnet surveillance under the misrepresentation that it's not a search until a human pulls the data. But a small on-device "E2E preserving" "safety" model that runs on your phone and snitches when illegal communication content is suspected. Edit: also consider centralized Room-641A-type surveillance when models summarize and/or flag all calls processed by public telephony
I hope this doesn't come to BCI.
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#650Earlier quoted context omitted.
This farmer needs a tote.
You probably haven't met a determined goat yet.
I’m concerned about the farmer being on the water without supervision when he’s concerned about how his imaginary wolf will get across.