Live data from Hacker News

The unbearable cheapness of open weight models

jamesoclaire.com

101–110 of 195 posts

Re: The unbearable cheapness of open weight models

#101
post #95

One issue I keep seeing with cost comparisons is that they compare API rates while a substantial fraction of users are on subscription plans. It's more expensive to use GLM 5.2 paying z.ai or Opencode Zen API rates than it is to use Opus on a subscription plan. Both of those providers offer subscriptions priced favorably relative to their API rates, but only in what are effectively trial sizes.

And that means either:

1. They overprice their APIs to make their subscriptions look reasonable

2. They burn money with their subscriptions

Re: The unbearable cheapness of open weight models

#102

With cache hit rates being effectively free, harnesses like Reasonix have let me do a month of work for less than 2 dollars. It's not even the subsidies making it cheap, American providers like Digital Ocean or Cloudflare host the same model with similar pricing.

Cloudflare's Deepseek V4 Pro prices are 4x more than Deepseek's for input and output tokens, and 100x more for cached input tokens, which is crucial for the tool uses of agents which cause multi-turn conversations.

Cache hit is less than a cent with Deepseek Flash and 3 cents with Cloudflare, it's free vs almost free. Where are you finding the statistics on Deepseek Pro? I don't see Cloudflare as a provider on openrouter for Pro, only flash.

Re: The unbearable cheapness of open weight models

#103
post #54

With cache hit rates being effectively free, harnesses like Reasonix have let me do a month of work for less than 2 dollars. It's not even the subsidies making it cheap, American providers like Digital Ocean or Cloudflare host the same model with similar pricing.

How does caching help here? How much repetition is there in queries?

Their blob explains it best, although this is a link to an older version design: https://github.com/esengine/DeepSeek-Reasonix/blob/v1/docs/A...

From my understanding if previous tokens are frozen and guaranteed to be immutable you can leverage that.

Re: The unbearable cheapness of open weight models

#104

It would not be surprising if GPT and Claude get cheaper too as inference gets cheaper. Two years ago, o1 was the strongest model and cost much more than Fable, while being nowhere near as smart as a Qwen 3.6 35B that you can now run on a DGX Spark without much trouble.

> It would not be surprising if GPT and Claude get cheaper too as inference gets cheaper

No because the biggest factor in their current price is VC subsidization which has likely peaked if OpenAI is now serving ads and Anthropic has increased their API pricing

Re: The unbearable cheapness of open weight models

#105

I don’t get it. So many here are saying open weight models will kill the frontier labs. But open source and similar have tried to beat private companies everywhere all the time, and people still buy the best products even if great open source alternatives are available. Why wouldn’t this be the case for AI too?

The closest example I can think of is using a proprietary hosted database versus a self hosted open source option, like Oracle vs Postgres. OpenAI and Anthropic are each individually privately valued over $1T and Oracle is currently valued at half that. They’re not worthless, but they’re severely overvalued.

Re: The unbearable cheapness of open weight models

#106
Deepseek's price looks unsustainable. Ant have said their operating margin is 70%. A leaner company could maybe raise that to 90%.

Most of the cost of supplying inference compute is depreciation of the GPUs. Maybe Deepseek is anticipating a 50 year life for theirs.

Re: The unbearable cheapness of open weight models

#107
post #23

Earlier quoted context omitted.

3) Buy all the RAM, increasing the barrier to entry to push back the tide a bit, in time for a juicy IPO.

4) Make it illegal to use anything but regulated models.

a: If making it illegal fails, make it a Federal procurement requirement to use regulated models. Come up with an audit standard that only fits regulated models. Watch the preference trickle down.

Re: The unbearable cheapness of open weight models

#108
post #37

Open weight and local hosting is far, far cheaper. In every respect. Even support is cheaper, over time. However, it's difficult to sell this to businesses who want contracts and KPIs, not staff and commitments. Regulated industries will favour the closed sources, either by choice or mandate. The interesting question is whether they will have better models, or worse models. History says they will receive a worse serv…

Cheaper until you factor in security and liability, which are going to get increasingly salient over time.

Re: The unbearable cheapness of open weight models

#109

One of the purposes of open weight models is to create a moat. If there were no open models available, I think we'd see much more and better models coming from Europe by now. Right now, any startup wanting to build and sell a model needs to be substantially better than the open models, which has become increasingly difficult and expensive.

Europe has Mistral.

You and readers may be interested in Europe 2031

1. https://europe2031.ai/

Re: The unbearable cheapness of open weight models

#110
post #95

One issue I keep seeing with cost comparisons is that they compare API rates while a substantial fraction of users are on subscription plans. It's more expensive to use GLM 5.2 paying z.ai or Opencode Zen API rates than it is to use Opus on a subscription plan. Both of those providers offer subscriptions priced favorably relative to their API rates, but only in what are effectively trial sizes.

Enterprise plans don't have the equivalent of the subsidized-usage-included Claude Max/ChatGPT Pro plans anymore. The revenue generated and total amount of tokens used by individuals is probably a tiny fraction of tokens billed at API pricing.
Post reply on HN