Live data from Hacker News

DeepSeek V4 Flash 0731

arcprize.org

71–80 of 479 posts

Re: DeepSeek V4 Flash 0731

#71
DeepSeek is my cheap and cheerful Chinese model of choice for API use. Has been for a while, but now it's Flash instead of Pro. Even cheaper, and now better then Pro. I feel like most of the major Chinese models are benchmaxxed, they have weird quirks every time I use them (Qwen 3.8 Max doesn't check its work and leaves stuff broken, doesn't write tests unless prompted, etc., Kimi ends up being quite expensive and rarely better than GPT Sol or Opus 5), while DeepSeek models seem to be generally as good as the benchmarks indicate: Not the best, but stronger across the board than any model within an order of magnitude of its price.

Re: DeepSeek V4 Flash 0731

#72

I’ve been refreshing hacker news constantly for a week now waiting for v4 pro, after they stated it would follow «soon». I have learnt «soon» is a matter of definition.

> been refreshing hacker news constantly for a week now waiting for v4 pro

https://reddit.com/r/DeepSeek is where the fellow F5ers are at.

Re: DeepSeek V4 Flash 0731

#74

I've been using it extensively since the release and the best summary I can give is that it's good enough to use it for (almost) everything and cheap enough that the cost are irrelevant. I'm running it in Oh My Pi with a second instance running as "advisor" and even with 5-6 active sessions (effectively 12 streams) I'm struggling to spend more than 5 bucks per day. OpenCode Go even has double limits temporarily so fo…

>Test coverage too low? Auto generate tests on CI for every pull-requests!

Terrible use-case.

Re: DeepSeek V4 Flash 0731

#75
post #4

Kimi K3 was an interesting model only a month ago, and now we're looking at the same performance for 1/20th of the price. Wild how fast this is advancing.

And now nobody seems interested in it because the price hasn't gone down it's still $3/$15 for all providers on openrouter because of some Kimi license https://openrouter.ai/moonshotai/kimi-k3#providers

Morph has it for a slight discount, apparently.

Uptime looks crap, though.

Re: DeepSeek V4 Flash 0731

#76
I love DeepSeek V4 Flash since the pre-0731, now even more. It is the first model that is truly too cheap to meter.

But I find it having a pretty significant problem with tool calling - no idea why, but tool calling with it is SLOW. As long as the model is reasoning, all good. But give it a bunch of tools and it becomes extremely slow.

Am I the only one experiencing this?

Re: DeepSeek V4 Flash 0731

#77
It's not frontier, but it's far past what we had at the beginning of the year. It's very usable. I get great instruction compliance, tool calling, and with a trivial workflows flow it has very good long-running performance as well.

Re: DeepSeek V4 Flash 0731

#80

Earlier quoted context omitted.

Dax (from Opencode) has tweeted that they can replicate or beat the price with rented GPUs. Deepseeks secret sauce is the incredibly cheap caching (magnitude cheaper than other providers). vLLM has recently released a similar approach. It's not as effective as what DeepSeek does but still an interesting development. I have no doubt that in due time other providers will match or perhaps even beat the current DeepSeek…

Eh, what are you guys even talking about? Deepseek is not cheapest provider as is, and it's MIT. So deepseek making it more expensive to use is just nonsense, they can only change their own pricing. It's the beauty of MIT license and open weights. If anything, these models are some of the safest in the world to use if you worry about a rug pull.

Cached input tokens are what drives most costs.
Post reply on HN