Live data from Hacker News

AMD acquires Taalas to boost inference performance by etching models in silicon

theregister.com

641–650 of 712 posts

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#641
I’m surprised I haven’t seen anyone mention video models yet. I don’t know how many fps 17000 tok/s translates to exactly but it’s gotta be a lot. Might make real time AI video possible.

Now that I think about it, real time AI video might be a clear case of “You scientists were so preoccupied with whether you could or not, you forgot to ask if you should”

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#642

Earlier quoted context omitted.

"seems like baking models into silicon is speed-running obsolescence" Now maybe. When models are flying passenger aircraft, other prerogatives will assert themselves. When a 50TB ROM means you can impulse purchase a ChatGPT 6.3 xhigh that runs on batteries, yet more use cases will be apparent.

Well, 50TB ROM Taalas HC1 style would be apparently a 400000b transistor system through a chip sized 2.5 meters on the side... :)

Yes, I know. This view is how these problems are always perceived, decade after decade, as our predecessors filled rooms with iron and silicon, unable to fathom that the equivalent capacity and power would be a portable device 20 years later. We're not at some end point in this process: the devices we have now will appear just a primitive in the years to come as a 10MB 5.25" Winchester drive appears to us now.

One of the underappreciated effects of the AI boom and associated money is that it has strongly reinvigorated R&D in hardware: it is clear that there is a real application for far greater density and lower power demand, and people are now pursuing this much harder than they had been. That will yield what it has always yielded; orders of magnitude jumps in capacity and performance.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#643
post #636

I see a lot of discourse about it being fast-to-deprecation. But I see it a different way personally. Modern LLMs are trying to do more with less. Focus on doing the right thing the first time. Even if we squeeze dumb LLMs, the significantly faster speed means quicker iterations. So a bad decision doesn't cost the time and inference costs that it cost before. It theoretically changes the scale of errant token spend.…

I agree with this take in a lot of ways. If you slash the token cost and increase speed for each token 1,000x, who cares if it takes even 20x as many tokens to achieve the goal?

And also, there are lots of tasks where models today are fine with doing. If you think of these things like appliances, who cares if it's not quite as powerful as the next generation? It was purchased to do a task, it still does that task very well. It feels like being in the 90s and asking "why buy a server today when they're going to be faster next year? Just keep renting mainframe time." Well maybe I just need a box to run our HR and payroll system, and this box manages to run it fine today.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#644
post #633

Earlier quoted context omitted.

Now the fun part, how will having an LLM in my washing machine help anything

You load it, tell it what's in and it sets the program, tells you what it set and why. You approve and off it goes

I do that today by turning a dial and pressing the start button.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#645
post #121

Given the fast churn of the models, how does it work out? Won’t the silicon etched model already be 1 or more versions behind by the time the silicon comes out. Though if it’s cheap enough, there certainly can be a market for cheaper model inferences.

also, as this scales, what would this mean for closed-weight hosted models? i imagine it's possible (but difficult) to re-derive model weights by de-lidding and inspecting the die... so will this only ever be used for open-weights models?

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#646
post #640

Earlier quoted context omitted.

Taalas exploits the low cardinality to store one 4-bit weight with one transistor. (They are using metal layer traces for the ROM, and connecting an access transistor to light up one of 16 options.) Their system is honestly very efficient for the weights, the problem is the KV-cache. That's why HC1 only supports such short context, they use SRAM for that and spend most of what's left of the die for it. The recent adv…

Do you really need a KV cache if inference is that fast though?

... Yes. Quadratic is really bad for large enough n, and you need that big context for useful work.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#647

Earlier quoted context omitted.

Slightly besides your point, but it's interesting how many here naturally ponder about how the current winner could or "should" keep winning, instead of how another company could become a competitor by doing the more clever thing the incumbent isn't thinking about.

It is not a “should”. At least not in the “we wish it were so” sense. It is more that there are multiple reasons why this idea (burning an LLM into silicone and deploying it into a device in people’s pockets) requires huge piles of cash and the kind of engineering chops only a few company posesses. Of course i would like it if a small upstart would do this, but it doesn’t seem likely as a posibility. They won’t have…

Yes, it's an interesting register (sorry for the claudism; blame lesswrong-weighted training) for the use of the word 'should'. I agree with your assessment and it is rarely articulated. Sometimes I think that the HN set is abused by big tech both from above on the employer side and the consumer usage side (all the T&C's, VC incentives and M&A taking away once-good-things). So they adopt the only sliver of agency-salving language available, like 'big company that I have no scope over should X'.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#648

I'm surprised there's not more discussion about potential inflection points here. When technology gets faster, it opens up whole new classes of UX that were hard to predict For example, faster internet didn't mean being able to view 100x as many HTML4 web pages. It brought SaaS, streaming media and interactivity. I'm not good at predicting, but some ideas: 1. All information gets augmented in real time with personali…

Right, instead of just having more LLM conversations as they get faster and cheaper we'll find products having AI that never had it before.

How about an actually smart thermostat that checks the weather and possibly makes decisions more like a human would i.e. tool calling, judgement, preference, history, personal plans.

Sure we have thermostats and you can configure rules and data sources, hook up Google calendar, etc but it has to be all predetermined and breaks as soon as anything stops working. AI could make this less brittle. AI agents are more flexible.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#649

Earlier quoted context omitted.

4. Pervasive , distributed dragnet surveillance under the misrepresentation that it's not a search until a human pulls the data. But a small on-device "E2E preserving" "safety" model that runs on your phone and snitches when illegal communication content is suspected. Edit: also consider centralized Room-641A-type surveillance when models summarize and/or flag all calls processed by public telephony

I hope this doesn't come to BCI.

BCI will only arrive once this style of monitoring can be enforced.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#650
post #488
post #353

Earlier quoted context omitted.

This farmer needs a tote.

You probably haven't met a determined goat yet.

Sometimes the goat will fill up on the tote and you can get the cabbage across, but you can’t count on it.

I’m concerned about the farmer being on the water without supervision when he’s concerned about how his imaginary wolf will get across.

Post reply on HN