In the three options OP presents, I wonder if there's a fourth: BYO model Customers give vendors metered access to their model. They can budget tokens per vendor. Vendors selling "AI products" can have a cleaner story and win on the margin. The first step to is to iron out a reasonable protocol, basically authorizing a, access token, and then the model providers (OpenAI, Anthropic, etc.) do the rate limiting. Theoret…
The current AI pricing was always going to go away
71–80 of 100 posts
Re: The current AI pricing was always going to go away
#72Earlier quoted context omitted.
Your concern about their business is that demand for their products is growing so stratospherically that they cannot meet that demand easily? I mean that's like an A+ scorecard for any business. Everyone in business would dream of such a scenario. That's called a luxury problem.
But each new customer is still losing money. As I said, subsidized growth only matters if you can recoup those subsidies afterwards - and that's what I'm not sure will be true. I think the idea of "all growth is good no matter the cost" has been taken to an extreme.
Re: The current AI pricing was always going to go away
#73Inference costs absolutely did fall. And even more so when looking at intelligence it buys you. eg compare say gpt 3.5 to latest deepseek. Both cheaper and more at more capable
Re: The current AI pricing was always going to go away
#74This is slightly more tasteful slop than average (I'm thinking probably Claude rather than ChatGPT?), but it's still 100% AI written: https://www.pangram.com/history/c55ab69b-e0a9-49a0-8056-2fcd...
Re: The current AI pricing was always going to go away
#75Guys, we are the in the mainframe era of AI. People in the 60's thought computing was expensive too and the idea of having a computer on every desk, nevermind every pocket, nevermind every single piece of electronics in the world basically seemed like a complete pipe dream. if you told someone in the 70's their toaster would have a supercomputer it in, they would think you were crazy. in 10 years your doorknob is goi…
The main issue with this reasoning is that the hardware substrate for AI and good old computing is the same. All governed by Moore's Law, what happened then seems extremely unlikely to happen again, the curve is a sigmoid and we're much closer to the flat end now.
Re: The current AI pricing was always going to go away
#76Earlier quoted context omitted.
I know it comes off as pedantic to point this out but: Those are open weight models not open source models. Closed weight models are the equivalent of SaaS. Open weight models are the equivalent of binary driver blobs or Windows software. We don't really have actual open source LLMs, which would need to publicly release their training data and technique so you could train a similar model yourself, or use their work a…
I know this is highly contested, but I'll try explaining it anyway, because I keep seeing this and it's ... wrong. Your comment is wrong both theoretically and practically. First, the theory. The idea that model weights are "binary driver blobs" is technically wrong. I don't know why this is so common on a technical site, but anyway. An LLM model consists of 3 main parts: The architecture, the inference code, and som…
Re: The current AI pricing was always going to go away
#77This is where open source models are important. The latest deepseek v4 pro model is 2-5x cheaper than Claude Sonnet 4.6. Cursor's Compose 2.5 that was just recently released is 6x cheaper than Sonnet. The state of the art models are going to get better and more expensive and smaller models are going to get cheaper. There will be a point where the intelligence of both the cheap and state of the art models are indistin…
They do matter in that oss researchers enable faster cross-pollination of good inferencing efficiency improvements to help the big boys adapt ideas from the community
Long-term local ai may matter more, but imo not there until models + hw get way better (1-2 years?) . Reasoning grade quality at speed is still $$$: we need fast opus, not slow sonnet.
Re: The current AI pricing was always going to go away
#78That is the question. I love using OpenCode with paid inference providers and seeing the cost of every little thing I do. On the other hand, right now I am flipping between Antigravity CLI and the two Antigravity apps burning Claude Opus tokens like crazy, knocking off a ton of work. Google must be losing money on me.
Re: The current AI pricing was always going to go away
#79Earlier quoted context omitted.
The labs have a perverse incentive to make things as expensive compute wise as possible. The only thing keeping this somewhat in check is competition, but it's intentionally being gatekept by locking up the supply of computing infrastructure. With 3 players it's pretty easy to collude even if indirectly. They can't burn trillions forever. Nvidia's 75% profit margins are not sustainable forever. Things will normalize,…
>The labs have a perverse incentive to make things as expensive compute wise as possible. The only thing keeping this somewhat in check is competition, but it's intentionally being gatekept by locking up the supply of computing infrastructure. With 3 players it's pretty easy to collude even if indirectly. By all accounts the AI capex boom is justified up by actual usage, rather than some nefarious plan for "locking u…
Re: The current AI pricing was always going to go away
#80This is where open source models are important. The latest deepseek v4 pro model is 2-5x cheaper than Claude Sonnet 4.6. Cursor's Compose 2.5 that was just recently released is 6x cheaper than Sonnet. The state of the art models are going to get better and more expensive and smaller models are going to get cheaper. There will be a point where the intelligence of both the cheap and state of the art models are indistin…
>The latest deepseek v4 pro model is 2-5x cheaper than Claude Sonnet 4.6. Cursor's Compose 2.5 that was just recently released is 6x cheaper than Sonnet. The only way you're running Deepseek V4 with comparable quality/performance is through OpenRouter, at which point you're still susceptible to being price gouged in the future, or by spending >$20k on hardware.