Earlier quoted context omitted.
There's only so far engineers can optimise the underlying transformer technique, which is and always has been doing all the heavy lifting in the recent ai boom. It's going to take another genius to move this forward. We might see improvements here and there but the magnitudes of the data and vram requirements I don't think will change significantly
I’ve read and heard from Semi Analysis and other best-in-class analysts that the amount of software optimizations possible up and down the stack is staggering… How do you explain that capabilities being equal, the cost per token is going down dramatically?
The beginning of scarcity in AI
231–239 of 239 posts
Re: The beginning of scarcity in AI
#232Earlier quoted context omitted.
I’ve read and heard from Semi Analysis and other best-in-class analysts that the amount of software optimizations possible up and down the stack is staggering… How do you explain that capabilities being equal, the cost per token is going down dramatically?
Optimizations, like I said. They'll never hack away the massive memory requirements however, or the pre training... Imagine the memory requirements without the pre training step....this is just part and parcel of the transformer architecture.
Re: The beginning of scarcity in AI
#233Earlier quoted context omitted.
If only there were some form of cheap, widely manufactured power generation technology that didn't use turbines... Are they really going to wait until 2030 to get more turbines rather than invest in solar?
I am clueless in this field, but solar seems to be unreliable and yield fraction of power required. Do you have a suggestion on something to read and learn more?
Amount of power: World solar power generation capacity is in the terawatts and rapidly increasing, there's no issue with its potential ceiling. As a bonus, it tends to work best on land that's useless for other purposes.
Re: The beginning of scarcity in AI
#234This is probably even the "fun" part of the whole picture. The purely dystopia starts when investment firms just silently grow bigger and bigger data centers like cancer. There will be no press releases, no papers, no chance anyone without billions will even know the details yet alone get access. One day we realise the worlds resources (maybe not as in the paperclip maximiser, but as in memory, energy, GPUs, water, l…
Re: The beginning of scarcity in AI
#235Earlier quoted context omitted.
Or simply by the fact that increasing production takes time? Any power plant takes years to build? Years, is like a lifetime for AI at this point...
> increasing production takes time? This is true of nearly everything (except money). I'm not sure of the point you are trying to make.
You can have all the money you want, if you want a gW of power added tomorrow, it won't happen.
Re: The beginning of scarcity in AI
#236Earlier quoted context omitted.
Or simply by the fact that increasing production takes time? Any power plant takes years to build? Years, is like a lifetime for AI at this point...
Not solar. China and to a lesser extent India are pumping out huge solar farms in months.
Re: The beginning of scarcity in AI
#237Earlier quoted context omitted.
> The companies that are entirely AI-dependent may need to raise prices dramatically as AI prices go up. It's not that clear. Sure, hardware prices are going up due to the extremely tight supply, but AI models are also improving quickly to the point where a cheap mid-level model today does what the frontier model did a year ago. For the very largest models, I think the latter effect dominates quite easily.
There's only so far engineers can optimise the underlying transformer technique, which is and always has been doing all the heavy lifting in the recent ai boom. It's going to take another genius to move this forward. We might see improvements here and there but the magnitudes of the data and vram requirements I don't think will change significantly
If you play around with the math, you quickly realize that even if we heavily quantize models down to INT4 to save memory, simply scaling the context window (which everyone wants now) immediately eats back whatever VRAM we just saved. The underlying math is extremely unforgiving without fundamentally changing the architecture.
Re: The beginning of scarcity in AI
#238We just had a realization during a demo call the other day: The companies that are entirely AI-dependent may need to raise prices dramatically as AI prices go up. Not being dependent on LLMs for your fundamental product’s value will be a major advantage, at least in pricing.
Re: The beginning of scarcity in AI
#239Earlier quoted context omitted.
> In 5 years consumer chips and model inference will be so good you won't need a server for SOTA. Naw man, you crazy. If you tell me that in 5 years, consumer chips will be so good that I can run GPT-5.4-level AI on my phone, I'd find that plausible (I buy cheap phones). If you're telling me that in 5 years we won't need _servers_ because our _phones and/or desktops_ will be powerful enough to run the biggest newest…
The thing is SOTA has a plateau. All LLMs work on the same principle: input goes in for training, reinforced by humans. There is only so much input (all recorded human knowledge), only so many human tweaks, that can produce only so much increased signal-to-noise in output. The machine can't read your mind, and there is no one truthful answer to most questions, so there will always be a limit on how accurate or correc…
The history of computing is full of predictions that consumer hardware would catch up to server-class capability in X years, and the answer has consistently been, consumer hardware catches up to _yesterday's_ server capability while server capability has moved on to new more mind-blowing paradigms which would not be possible on consumer hardware for another half-decade or more.
I'm sure that specific scaling trajectories will hit specific ceilings, such that in specific ways, one can make the argument that (for example) today's iphone performs at parity with today's servers. In 5 minutes I can spin up the same Postgres or Mongo DB that the largest companies on earth use server-side, though I can't support anywhere near the same data & traffic volume. But parity along specific technical aspects is a very different matter from the broad prediction of "you won't need a server for SOTA".
To step back to the bigger context -- your original point seems more along the lines of "we're obviously in an unsustainable bubble, and the rapid progress in on-device AI will further exacerbate the embarrassing collapse of all these overhyped AI companies". I strongly agree with you. But I think that's likely _and also_ firmly predict that the technical SOTA of 2031 (and 2041, if we make it there), in nearly every imaginable aspect including language-capable AI, will be vastly more capable than what you can run in your pocket.