Live data from Hacker News

The current AI pricing was always going to go away

arnon.dk

11–20 of 100 posts

Re: The current AI pricing was always going to go away

#11

This is where open source models are important. The latest deepseek v4 pro model is 2-5x cheaper than Claude Sonnet 4.6. Cursor's Compose 2.5 that was just recently released is 6x cheaper than Sonnet. The state of the art models are going to get better and more expensive and smaller models are going to get cheaper. There will be a point where the intelligence of both the cheap and state of the art models are indistin…

Deepseek V4 Flash is far cheaper still, and a better model to compare to Sonnet 4.6. I'm finding it a reliable workhorse.

Re: The current AI pricing was always going to go away

#12
Insofar as I can tell, inference is on a certain path toward becoming "free". The models are now extremely powerful on high-end consumer hardware, and the efficiency trend seems likely to continue.

Here is a recent non-rigorous benchmark I ran against a bunch of models. Qwen3.6 35B A3B fine-tuned with opus data runs plenty fast on my local machine and produce outstanding results - easily in the top 5, comparable to GPT 5.5 Pro (which is $180/mtok).

https://gistpreview.github.io/?31d66ef69e4aed3efae1aec69d86c...

I've predicted for years now that the industry will head down the path of the virus scanning vendors: selling subscriptions to be able to download the latest versions of models. I simply don't see how any other business model is remotely viable, except at the very highest end of inference or video gen.

Re: The current AI pricing was always going to go away

#13
post #10
post #5

This is slightly more tasteful slop than average (I'm thinking probably Claude rather than ChatGPT?), but it's still 100% AI written: https://www.pangram.com/history/c55ab69b-e0a9-49a0-8056-2fcd...

This... is not a reliable AI detection method at all.

Pangram is highly reliable.

Re: The current AI pricing was always going to go away

#14

This is where open source models are important. The latest deepseek v4 pro model is 2-5x cheaper than Claude Sonnet 4.6. Cursor's Compose 2.5 that was just recently released is 6x cheaper than Sonnet. The state of the art models are going to get better and more expensive and smaller models are going to get cheaper. There will be a point where the intelligence of both the cheap and state of the art models are indistin…

> The state of the art models are going to get better and more expensive and smaller models are going to get cheaper.

Why do you think this will be true?

Right now I see the major US labs betting on gaining an advantage from having way more compute, and I see Chinese labs competing with one another in a resource-scarce environment, so they place much more emphasis on compute-efficiency.

But the supply chains that feed into the massive data center growth in the US are strained; there are energy, memory, and logistical bottlenecks to name a few.

In the medium-long run, compute capacity will not grow exponentially forever. Somehow it has for decades, but there can be no infinite exponential growth, and that point may be when the planet really starts to cook itself.

Maybe the US labs will become more compute-constrained, and then have to compete on efficiency.

Or maybe things change fundamentally in some other way I'm not thinking of.

Re: The current AI pricing was always going to go away

#16
post #10
post #5

This is slightly more tasteful slop than average (I'm thinking probably Claude rather than ChatGPT?), but it's still 100% AI written: https://www.pangram.com/history/c55ab69b-e0a9-49a0-8056-2fcd...

This... is not a reliable AI detection method at all.

You are incorrect. There, now we've both made unsupported assertions. Care to provide any evidence for your position?

For what it's worth, when I provide a Pangram link it's because I can already tell something is AI and I'm attempting to provide objective third-party confirmation so the conversation doesn't just degrade into me asserting that I have superior taste to you.

Re: The current AI pricing was always going to go away

#18

I wonder how much of Uber blowing their AI budget and MSFT pulling their claude code licenses can be attributed to "tokenmaxxing". When Meta announced token leaderboards and other followed, I could see this being the logical conclusion. That whole trend is so dumb because it leads to this. Company announces they will measure developer performance by how many tokens they burn and constantly talks about how the best de…

> I personally use my OpenAI subscription pretty heavily, 2-3 agents running practically all day on various tasks but I never even get close to running into limits

Same. But if I was working for an organization that measured token usage, you can bet I would be doing things like creating a cron job that uses claude to create a customized bespoke report update of the current status of all my open assigned tickets and message that to myself 4 times a day... token burn for zero purpose whatsoever.

Re: The current AI pricing was always going to go away

#19
You are comparing two different model. It's like saying roadster is more expensive than model S. No model pricing actually increased, and I am using GPT-4o in the same price as it was before.

You can see price vs performance in artificial analysis and the the pareto optimal is all just 6 months old model.

Re: The current AI pricing was always going to go away

#20

This is where open source models are important. The latest deepseek v4 pro model is 2-5x cheaper than Claude Sonnet 4.6. Cursor's Compose 2.5 that was just recently released is 6x cheaper than Sonnet. The state of the art models are going to get better and more expensive and smaller models are going to get cheaper. There will be a point where the intelligence of both the cheap and state of the art models are indistin…

> The state of the art models are going to get better and more expensive and smaller models are going to get cheaper. Why do you think this will be true? Right now I see the major US labs betting on gaining an advantage from having way more compute, and I see Chinese labs competing with one another in a resource-scarce environment, so they place much more emphasis on compute-efficiency. But the supply chains that fee…

The labs have a perverse incentive to make things as expensive compute wise as possible. The only thing keeping this somewhat in check is competition, but it's intentionally being gatekept by locking up the supply of computing infrastructure. With 3 players it's pretty easy to collude even if indirectly. They can't burn trillions forever. Nvidia's 75% profit margins are not sustainable forever.

Things will normalize, but it will take time.

Post reply on HN