Tokens will surely become commodities, but models don't need to. The point is that electricity becomes a commodity, rather than, say, nuclear reactors or gas turbines also need to be commoditized. Contrary to most people in the software-minded circle, I am not very excited about running local LLMs. Most of what I have seen involves selling wrappers that call APIs on top of an LLM AI, similar to how traditional SaaS h…
Ways to think about token pricing
21–26 of 26 posts
Re: Ways to think about token pricing
#22I know there's a lot of reasons to think that everyone will just use AI inference in the cloud, but I think if everyone had access to a dgx gb400 class machine with 512gb of hbm4 vram and 1tb of lpddr8x a lot of people are going to be running finetunes of models locally. Like the dgx gb300 is $94k now, but i bet this class of machine will come down to $20kish in the next 2-3 years.
if all the competition pushes down margins on tokens to 10-20%, i dont see how the inherent scale advantages of cloud inference wont be way more to account for the 20% cheaper tokens youd be getting running locally. i dont see how local will ever be more economical
Re: Ways to think about token pricing
#23Earlier quoted context omitted.
sure, if there werent capitalists digging their moat. in 1 year local inferwnce doubled cost from ram alone because cloud vendors cornered markets on their circular cash flow expectations.
There’s a shitload of money to be made in RAM fabrication. Costs will come down a ton.
Re: Ways to think about token pricing
#24Re: Ways to think about token pricing
#25Re: Ways to think about token pricing
#26anyone actually tracking cost per completed task rather than cost per call ? it is quite a difficult task to define completion as it is a qualitative aspect and can vary from every individual.