Live data from Hacker News

DeepSeek V4 Flash 0731

arcprize.org

31–40 of 477 posts

Re: DeepSeek V4 Flash 0731

#32
post #15

Earlier quoted context omitted.

But DeepSeek now has a warning they’re going to sharply increase their API pricing sometime in the future.

Dax (from Opencode) has tweeted that they can replicate or beat the price with rented GPUs. Deepseeks secret sauce is the incredibly cheap caching (magnitude cheaper than other providers). vLLM has recently released a similar approach. It's not as effective as what DeepSeek does but still an interesting development. I have no doubt that in due time other providers will match or perhaps even beat the current DeepSeek…

Eh, what are you guys even talking about? Deepseek is not cheapest provider as is, and it's MIT. So deepseek making it more expensive to use is just nonsense, they can only change their own pricing. It's the beauty of MIT license and open weights. If anything, these models are some of the safest in the world to use if you worry about a rug pull.

Re: DeepSeek V4 Flash 0731

#34
post #15

I've been using it extensively since the release and the best summary I can give is that it's good enough to use it for (almost) everything and cheap enough that the cost are irrelevant. I'm running it in Oh My Pi with a second instance running as "advisor" and even with 5-6 active sessions (effectively 12 streams) I'm struggling to spend more than 5 bucks per day. OpenCode Go even has double limits temporarily so fo…

But DeepSeek now has a warning they’re going to sharply increase their API pricing sometime in the future.

Yes, there is warning, but also there are many providers on OpenRouter[0], hosting open weight model with similar pricing. The question is Will they go up as well?

[0] https://openrouter.ai/deepseek/deepseek-v4-flash-0731#provid...

Re: DeepSeek V4 Flash 0731

#35
post #6
post #4

Kimi K3 was an interesting model only a month ago, and now we're looking at the same performance for 1/20th of the price. Wild how fast this is advancing.

Yeah either the benchmark isn't very useful anymore or V4 Flash is a really, really good model.

In my use, DeepSeek v4 Flash (which replaced the quite excellent MiniMax M3) lags behind GLM 5.2 & Muse Spark 1.2 (let alone Kimi K3). Also, K3 is a much bigger multi-modal model, while Flash is text-only and likely optimised for coding tasks.

Re: DeepSeek V4 Flash 0731

#36
post #4

Kimi K3 was an interesting model only a month ago, and now we're looking at the same performance for 1/20th of the price. Wild how fast this is advancing.

And now nobody seems interested in it because the price hasn't gone down

it's still $3/$15 for all providers on openrouter

because of some Kimi license

https://openrouter.ai/moonshotai/kimi-k3#providers

Re: DeepSeek V4 Flash 0731

#37
post #2

results comparable to gpt 5.6 luna but cheaper promising!

Is it still cheaper than Luna if using an OpenAI subscription? My gut is no, but I have not done the math.

Everything is cheaper if using a subscription, but some applications require API usage.

Re: DeepSeek V4 Flash 0731

#39

Earlier quoted context omitted.

Dax (from Opencode) has tweeted that they can replicate or beat the price with rented GPUs. Deepseeks secret sauce is the incredibly cheap caching (magnitude cheaper than other providers). vLLM has recently released a similar approach. It's not as effective as what DeepSeek does but still an interesting development. I have no doubt that in due time other providers will match or perhaps even beat the current DeepSeek…

Eh, what are you guys even talking about? Deepseek is not cheapest provider as is, and it's MIT. So deepseek making it more expensive to use is just nonsense, they can only change their own pricing. It's the beauty of MIT license and open weights. If anything, these models are some of the safest in the world to use if you worry about a rug pull.

90%+ cache hit rate is common, and so you'll see on places like openrouter that Deepseek cache cost is indeed a magnitude cheaper than the rest.

Re: DeepSeek V4 Flash 0731

#40
post #15

Earlier quoted context omitted.

But DeepSeek now has a warning they’re going to sharply increase their API pricing sometime in the future.

Dax (from Opencode) has tweeted that they can replicate or beat the price with rented GPUs. Deepseeks secret sauce is the incredibly cheap caching (magnitude cheaper than other providers). vLLM has recently released a similar approach. It's not as effective as what DeepSeek does but still an interesting development. I have no doubt that in due time other providers will match or perhaps even beat the current DeepSeek…

any link to this caching tech?
Post reply on HN