DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis
91–100 of 342 posts
Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis
#92Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis
#93The weights were just released a few minutes ago: https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731
Can't wait for the DwarfStar quants - I have been using DeepSeek v4 flash (preview) as my main coding agent for months now (running on my 128gb mbp) - it seems this model outperforms GLM 5.2 on nearly every metric. Thanks for sharing the news, I was refreshing huggingface but gave up thinking it likely would take some more time.
Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis
#94[flagged]
Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis
#95Earlier quoted context omitted.
I’ve been using v4 flash for an app I’m building [1] and it’s amazing how cost effective and good it is coming from having always used gpt, opus and sonnet models. It’s so cost effective I can offer a generous free tier since my goal isn’t to make money with it. [1] https://trysojourn.app
Is the difference between this and a frontier model that the scripture is guaranteed to be real? I'm on a team that develops a Bible study app, and we're all relatively content with how the basic models converse regarding scripture. Even as far back as GPT-4 was excellent. They occasionally have minor hallucinations (a dealbreaker for a production app), but they do an excellent job with theology and Bible scholarship…
flash is suitable only for a toy apps, not for production environments :)
Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis
#96Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis
#97Earlier quoted context omitted.
How exactly will they ban them?
By making companies using them "toxic" to touch. For example: no government contract to any company who uses even one vendor in it's entire chain of dependencies, who uses such open models. They can extend this further by laying more conditions, such as: any company dealing in this-this field can only use models "officially" approved as "safe". Rest you can guess how easy it would be to get that "safe" rating for suc…
Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis
#98here it is on openrouter https://openrouter.ai/deepseek/deepseek-v4-flash-0731
Why do the cache hit rates seem to vary so much between harnesses?
I use pi, which is very minimalist, and I get a hit rate of ~99%. Paying like $1 a day for Flash. Yet, the hit rate mentioned on OpenRouter is only ~79%.
Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis
#99Already beat Luna on price/task, by about 2x: https://artificialanalysis.ai/models/deepseek-v4-flash?intel...
Maybe I'm reading that incorrectly, but it seems to me the cost is on the X-axis. First, your direct comparison, Deepseek V4 Flash 0731 (max effort) $0.03 (rounded up) per task @ index 50. OpenAI Luna: * high effort $0.03 (rounded down) @ index 46 * xhigh effort $0.04 @ index 49 * max effort $0.07 @ index 51 So I would say a fair statement would be "OpenAI Luna between 2x and 3x the price of Deepseek Flash, what you…
uhh openai is dark gray: `rgb(31, 31, 31)` and i'm pretty sure it always has been?
Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis
#100It’s exciting that a model scoring this high is dirt cheap. It’s also so inefficient, when they release the full performance numbers it’s not going to be good. One example, it takes about 3.6x more tokens to finish the same work as Gemini Flash 3.6.
Using tokens to evaluate models is an outdated approach. Cost per task is what matters. Not all tokens are created equal