Earlier quoted context omitted.
Also, there is zero reason to think that the big labs did not have anything similar to TurboQuant for a long time already. The recent blog post from Google announcing TurboQuant does not change anything regarding RAM planning for the big labs. TurboQuant itself is already a year old! So even smaller labs have probably seen and implemented it.
TurboQuant has a specific benefit by compressing the KV cache at a negligible cost to quality. That mainly means that the context lengths can go up in models for the same amount of memory, however the KV cache only accounts for something like 20% of the overall model size, and this will not dramatically decrease memory demands in the way that some of the more sensationalist reporting has stated.
How the AI Bubble Bursts
321–330 of 557 posts
Re: How the AI Bubble Bursts
#322Okay lets suppose all those companies are profitable if training would stop today. What if token demand is shrinking ? I think big parts of the current demand is artificially build by e.g. FOMO and marketing without real value generated by them. There is no indication in economic data about some productivity boom resulting from AI usage. Next thing is Energy costs - that will soon eat into profitability too. I don't…
I don't think token demand will shrink because we're still just learning how to use it, demand will skyrocket. The problem is what price we'll be willing to pay for it, specially if competition keeps soaring.
Re: How the AI Bubble Bursts
#323> RAM prices are crashing because new models won’t need as much Reality begs to differ [0] and following the link for that text goes to an article [1] where they talk about Google's TurboQuant which supposedly will lower the RAM requirements. Now if that means RAM prices come down (as speculated, not reported on, in the link) or the AI companies just do more things with their extra ram is yet to be determined. The fa…
> Reality begs to differ [0] and following the link for that text goes to an article [1] where they talk about Google's TurboQuant which supposedly will lower the RAM requirements. Now if that means RAM prices come down (as speculated, not reported on, in the link) or the AI companies just do more things with their extra ram is yet to be determined. The fact this article links there with text "RAM prices are crashing" throws the entire rest of the article into doubt for me.
I find it fascinating how extremely reactive things have become. One research paper which, to my knowledge, hasn't been externally replicated yet, nor implemented, generate tons of hyperbolic article, tweets and such, and do actually manage to move the market at least temporarily. Not just this, but a simple message in full caps lock by the president of the U.S who is in the habit of lying through is teeth constantly, and the same thing happens. It's like there is a big bubble that threw any form of critical thinking out of the window and is in a hurry to react to anything even if it is not even remotely believable. Now I understand why it happens, there is a lot of money that can be made by capitalizing on FOMO, either by driving traffic to their website, socials, etc, or by simply insider trading (which feels like it has been legalized these days). But I still find it incredible the proportion it started to take.
Re: How the AI Bubble Bursts
#324> RAM prices are crashing because new models won’t need as much Reality begs to differ [0] and following the link for that text goes to an article [1] where they talk about Google's TurboQuant which supposedly will lower the RAM requirements. Now if that means RAM prices come down (as speculated, not reported on, in the link) or the AI companies just do more things with their extra ram is yet to be determined. The fa…
> > RAM prices are crashing because new models won’t need as much > Reality begs to differ [0] and following the link for that text goes to an article [1] where they talk about Google's TurboQuant which supposedly will lower the RAM requirements. Now if that means RAM prices come down (as speculated, not reported on, in the link) or the AI companies just do more things with their extra ram is yet to be determined. Th…
Re: How the AI Bubble Bursts
#325> RAM prices are crashing because new models won’t need as much Reality begs to differ [0] and following the link for that text goes to an article [1] where they talk about Google's TurboQuant which supposedly will lower the RAM requirements. Now if that means RAM prices come down (as speculated, not reported on, in the link) or the AI companies just do more things with their extra ram is yet to be determined. The fa…
> > RAM prices are crashing because new models won’t need as much > Reality begs to differ [0] and following the link for that text goes to an article [1] where they talk about Google's TurboQuant which supposedly will lower the RAM requirements. Now if that means RAM prices come down (as speculated, not reported on, in the link) or the AI companies just do more things with their extra ram is yet to be determined. Th…
Re: How the AI Bubble Bursts
#326Re: How the AI Bubble Bursts
#327Earlier quoted context omitted.
> Now if that means RAM prices come down (as speculated, not reported on, in the link) or the AI companies just do more things with their extra ram is yet to be determined. I think it is determined: https://en.wikipedia.org/wiki/Jevons_paradox
Jevons paradox only applies if demand hasnt already been saturated. The fact that public LLM usage is leveling off at a price of $0 and Jensen "we make the shovels in this gold rush" Huang is rather desperately claiming that you need to spend $250k/year in tokens to be taken seriously suggests that demand saturation may not be that far off. Whether Jevons' Paradox applies to software engineers I think is another open…
The ceiling of token use when everyone has something akin to OpenClaw just running as a background process on their phone is way higher than there’s supply for right now. Jevons paradox is still in full force.
Re: How the AI Bubble Bursts
#328Earlier quoted context omitted.
Demand of tokens is absolutely skyrocketing. And unlike the traditional "this will replace humans right away", I think what this introduce is a lot of incentive to spread those token in places where there was never any incentive to hire a software engineer for previously. In turn, that will drive a lot of business activity in those area that will potentially fail given the current quality of the output. This feels li…
Tulips sales also skyrocketed. Seriously, what value are tokens providing other than justifying layoffs. Concretely. Today. Not in the speculating scenario that cardiologist could be replaced with models. We see this new trend of agentic coding, again a promise software will be written that way going forward, despite the number of fiasco already experienced when trusting a model turned bad. The use case may provide v…
Re: How the AI Bubble Bursts
#329Earlier quoted context omitted.
That NeoLiberal shift did not take place in a vacuum. It was a product of the world around it. It absolutely was caused by tech. If we — those with the power to build the productivity creators — took a stand and said "we refuse to create tech for the interests of the few" it would have never happened. But, instead, we welcomed it and are responsible for it.
The corollary of “if we took a stand” is that Capital took a stand and collectively undid a lot of the gains of the post-WWII social democratic order. So no. It wasn’t caused by tech beyond the uninteresting factors like modern society being complex and, of course, that tech developments influence things (pretty much all things).
The benefactor of those gains was also entirely decided by those who created the tech. We could have given use of that tech to everyone. In some cases we actually did (e.g. open source), but in most cases we gained (at least partial) ownership of the capital so it was in our best personal economic interest to keep it for ourselves and our close friends.
Re: How the AI Bubble Bursts
#330Earlier quoted context omitted.
> The cost to serve tokens is absolutely profitable today and that’s been true for at least a year. > For the data center build outs, demand for tokens is still exceeding supply. Can you provide any numbers for this?
Check the token prices for open weight LLMs at various independent inference providers. That gives you a very good estimate of "how much can you serve the tokens of a model of the size N for while making a profit". Now, keep in mind: Kimi K2.5 is 1T MoE. Today's frontier LLMs are in the 1T to 5T range, also MoE. Make an estimate. Compare that estimate with the actual frontier lab prices.
In the current volatile environment, the API prices are more of a baseline where we can assume it can't be much cheaper to operate these models.