Is there scope to implement ternary models using this approach to minimise die area of the model parameters?
No gain in having less than FP4, so.
721–730 of 736 posts
Is there scope to implement ternary models using this approach to minimise die area of the model parameters?
No gain in having less than FP4, so.
Earlier quoted context omitted.
Yes, I know. This view is how these problems are always perceived, decade after decade, as our predecessors filled rooms with iron and silicon, unable to fathom that the equivalent capacity and power would be a portable device 20 years later. We're not at some end point in this process: the devices we have now will appear just a primitive in the years to come as a 10MB 5.25" Winchester drive appears to us now. One of…
> decade after decade That can only work when there is physical capacity for improvement though. > underappreciated effects of the AI boom and associated money is that it has strongly reinvigorated R&D in hardware Yes, absolutely: but the point at this stage is more about finding new possibilities in hardware architecture than the improvement of what we had. So > * That will yield what it has always yielded[:] orders…
There are great opportunities for advancement. Both in the physical hardware and in how and where it's deployed and powered.
Consider this, as only one point: there hasn't really been a demand for advancements in ROM. RAM has been scaling at approximately Moore's law rate, and nonvolatile R/W storage has been sedately scaling, but there hasn't been a use case for really dense, high performance ROM. Now there is. ROM used to be a big deal in computing and media (cartridges, optical disks, etc.,) but that tapered off long ago; volatile and R/W storage was sufficient and convenient for the time, and the inference model use case, where dense, high speed ROM can have extremely high value, didn't exist.
Now there is a use case, and industry is thinking about something they haven't cared about in a long time. Current fabrication nodes, stacked in the third dimension à la NAND flash, could produce staggeringly dense, fast and low power ROM. That's why AMD snatched up Taalas: they're thinking about an aspect of the future that has been (reasonably) neglected.
Earlier quoted context omitted.
> Is there any LLM from exactly one year ago that would be worth running? Bad perspective: consider the correction: "when are thresholds of sought quality reached"? Hence: not "is there a 10yo from last year that could compete with the current 13yo", but "will there be a 30(?)yo from last year that could compete with the current 33(?)yo" ('(?)': the scale of yearly growth in the future is uncertain).
It's not just about it "being smart enough". It's about there being actual user demand when it needs to compete with the shiny new model. A 10 year old iPhone is probably good enough, but is there demand for it? In a vacuum a 10 year old iPhone is good, but why would you pick it if you can have a current one for a reasonable price?
So, when the models will be "good enough", you will probably use one as the "daily driver" for a long time for consolidated workflows (some of them enabled by the staggering collateral advantages of specialized hardware and obvious advantages of local hardware), and occasionally use other available models for exceptional tasks, and upgrade only when definitely advantageous - like normal goods.
Earlier quoted context omitted.
> decade after decade That can only work when there is physical capacity for improvement though. > underappreciated effects of the AI boom and associated money is that it has strongly reinvigorated R&D in hardware Yes, absolutely: but the point at this stage is more about finding new possibilities in hardware architecture than the improvement of what we had. So > * That will yield what it has always yielded[:] orders…
> That can only work when there is physical capacity for improvement though. There are great opportunities for advancement. Both in the physical hardware and in how and where it's deployed and powered. Consider this, as only one point: there hasn't really been a demand for advancements in ROM. RAM has been scaling at approximately Moore's law rate, and nonvolatile R/W storage has been sedately scaling, but there hasn…
Clearly there are possibilities, some of them proven (proof-of-concept, in-production etc.) - but taking for granted "Moore's law" like spaces for them may not be founded on what we know at this stage.
I see a lot of discourse about it being fast-to-deprecation. But I see it a different way personally. Modern LLMs are trying to do more with less. Focus on doing the right thing the first time. Even if we squeeze dumb LLMs, the significantly faster speed means quicker iterations. So a bad decision doesn't cost the time and inference costs that it cost before. It theoretically changes the scale of errant token spend.…
Sounds to me like you would hire 20 barely-paid interns instead of 2 competent programmers.
I think what will happen is what happened to something like 4K video decoding before where it ends up in silicon costing almost nothing to run extremely fast on device. "Good enough" LLM functionality (for the use case) will be on-die or on-chip for cars, appliances, etc. This will provide speeds of chatjimmy at a battery-level power consumption. Probably this will also happen for software engineering. Some usb-power…
I think this combined with a bit of memory and something like the “high-bandwidth flash” they just announced (if it works out), could be an interesting thing for some resident (burned) experts + active moe streamed from HBF. I expect one day having small lower power demand drive sized devices with proprietary burned-in models that are quite fast running on-device in robotics and such. Commoditizing LLMs via burned an…
Earlier quoted context omitted.
I'd gladly pay for a Claude Opus 4.6 Thinking High in silicon and use it for 1-2 years. It's good enough for many coding tasks.
It costs something like $300,000 for the hardware to run a model of that size. You'd pay that for a single model for 1-2 years? Not even the AI companies can justify that kind of spend which is why they keep extending the expected lifespan on their hardware in the accounting.
Earlier quoted context omitted.
> That can only work when there is physical capacity for improvement though. There are great opportunities for advancement. Both in the physical hardware and in how and where it's deployed and powered. Consider this, as only one point: there hasn't really been a demand for advancements in ROM. RAM has been scaling at approximately Moore's law rate, and nonvolatile R/W storage has been sedately scaling, but there hasn…
Sure, but actually, our current need is not really for "ROM": it is for "CiM", compute-in-memory - we want to minimize the data movement bottlenecks. That some implementations could be read-only is actually a disadvantage. Clearly there are possibilities, some of them proven (proof-of-concept, in-production etc.) - but taking for granted "Moore's law" like spaces for them may not be founded on what we know at this st…