Live data from Hacker News

AMD acquires Taalas to boost inference performance by etching models in silicon

theregister.com

721–727 of 727 posts

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#721
post #570

Is there scope to implement ternary models using this approach to minimise die area of the model parameters?

To the best of my understanding it would make no sense - they are already using (I have to investigate how they achieved it technically) a single transistor to store the "weight", and they manage to encode an FP4 in it (through a table).

No gain in having less than FP4, so.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#722

Earlier quoted context omitted.

Chip pops out like a gameboy cartridge. AI not working? Blow on it and jam it back in

Agree. Price is the question.

Depends on how smart you want your MegaMan to be

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#723

Earlier quoted context omitted.

Yes, I know. This view is how these problems are always perceived, decade after decade, as our predecessors filled rooms with iron and silicon, unable to fathom that the equivalent capacity and power would be a portable device 20 years later. We're not at some end point in this process: the devices we have now will appear just a primitive in the years to come as a 10MB 5.25" Winchester drive appears to us now. One of…

> decade after decade That can only work when there is physical capacity for improvement though. > underappreciated effects of the AI boom and associated money is that it has strongly reinvigorated R&D in hardware Yes, absolutely: but the point at this stage is more about finding new possibilities in hardware architecture than the improvement of what we had. So > * That will yield what it has always yielded[:] orders…

> That can only work when there is physical capacity for improvement though.

There are great opportunities for advancement. Both in the physical hardware and in how and where it's deployed and powered.

Consider this, as only one point: there hasn't really been a demand for advancements in ROM. RAM has been scaling at approximately Moore's law rate, and nonvolatile R/W storage has been sedately scaling, but there hasn't been a use case for really dense, high performance ROM. Now there is. ROM used to be a big deal in computing and media (cartridges, optical disks, etc.,) but that tapered off long ago; volatile and R/W storage was sufficient and convenient for the time, and the inference model use case, where dense, high speed ROM can have extremely high value, didn't exist.

Now there is a use case, and industry is thinking about something they haven't cared about in a long time. Current fabrication nodes, stacked in the third dimension à la NAND flash, could produce staggeringly dense, fast and low power ROM. That's why AMD snatched up Taalas: they're thinking about an aspect of the future that has been (reasonably) neglected.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#724

Earlier quoted context omitted.

> Is there any LLM from exactly one year ago that would be worth running? Bad perspective: consider the correction: "when are thresholds of sought quality reached"? Hence: not "is there a 10yo from last year that could compete with the current 13yo", but "will there be a 30(?)yo from last year that could compete with the current 33(?)yo" ('(?)': the scale of yearly growth in the future is uncertain).

It's not just about it "being smart enough". It's about there being actual user demand when it needs to compete with the shiny new model. A 10 year old iPhone is probably good enough, but is there demand for it? In a vacuum a 10 year old iPhone is good, but why would you pick it if you can have a current one for a reasonable price?

If you have a 2 years old item working and the new 2 months old item offers little advantage over the old one, and its cost is nonzero, an important part of the aggregate demand will stick with the old one...

So, when the models will be "good enough", you will probably use one as the "daily driver" for a long time for consolidated workflows (some of them enabled by the staggering collateral advantages of specialized hardware and obvious advantages of local hardware), and occasionally use other available models for exceptional tasks, and upgrade only when definitely advantageous - like normal goods.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#725

Earlier quoted context omitted.

> decade after decade That can only work when there is physical capacity for improvement though. > underappreciated effects of the AI boom and associated money is that it has strongly reinvigorated R&D in hardware Yes, absolutely: but the point at this stage is more about finding new possibilities in hardware architecture than the improvement of what we had. So > * That will yield what it has always yielded[:] orders…

> That can only work when there is physical capacity for improvement though. There are great opportunities for advancement. Both in the physical hardware and in how and where it's deployed and powered. Consider this, as only one point: there hasn't really been a demand for advancements in ROM. RAM has been scaling at approximately Moore's law rate, and nonvolatile R/W storage has been sedately scaling, but there hasn…

Sure, but actually, our current need is not really for "ROM": it is for "CiM", compute-in-memory - we want to minimize the data movement bottlenecks. That some implementations could be read-only is actually a disadvantage.

Clearly there are possibilities, some of them proven (proof-of-concept, in-production etc.) - but taking for granted "Moore's law" like spaces for them may not be founded on what we know at this stage.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#727
post #655
post #636

I see a lot of discourse about it being fast-to-deprecation. But I see it a different way personally. Modern LLMs are trying to do more with less. Focus on doing the right thing the first time. Even if we squeeze dumb LLMs, the significantly faster speed means quicker iterations. So a bad decision doesn't cost the time and inference costs that it cost before. It theoretically changes the scale of errant token spend.…

Sounds to me like you would hire 20 barely-paid interns instead of 2 competent programmers.

Weird conclusion to draw from my statement. Feels intentionally hostile of an assessment about me, but I digress.
Post reply on HN