The weights were just released a few minutes ago: https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731
Can't wait for the DwarfStar quants - I have been using DeepSeek v4 flash (preview) as my main coding agent for months now (running on my 128gb mbp) - it seems this model outperforms GLM 5.2 on nearly every metric. Thanks for sharing the news, I was refreshing huggingface but gave up thinking it likely would take some more time.
DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis
161–170 of 342 posts
Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis
#162Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis
#163Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis
#164Does it already know the answer to what happen at Tiananmen Square? Or still avoiding it?
How often are you asking an LLM about this? You know you can just google it, right? I have to admit it rarely comes up in the coding tasks I usually give to LLMs.
Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis
#165So, are they planning to announce an optimized coding agent harness as well ? DSv4 flash is a fantastic model, and my daily driver. With reasonix or pi, I can code all day long and pay a few pennies for it. No token anxiety. Whereas the same model with fireworks/openrouter, with zdr thrown in, token costs ratchet up with no explanation. Likely that the model is subsidized for gathering usage data. I am waiting for the day I can run this locally.
Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis
#166Earlier quoted context omitted.
The thing is, you don’t need to actually block usage to make something illegal. You make it so toxic that company wants to be seen publicly using open models
Companies would secretly use self-hosted models internally because it would give them a enormous cost advantage.
Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis
#167Earlier quoted context omitted.
Indeed. My fellow software engineers keep complaining about using up all their Claude tokens within an hour... Whilst I'll be rocking DS flash for the entire day. Sure it gets a few things wrong here and there, but that's when you pull out the Claude models or whatever for those tricky tasks.
The problem is picking between models. I do not want to spend my time switching models and trying to decipher which model should be used for what. Maybe that's just a me problem that I need to figure out.
Of course you can get more bang for your buck by being more deliberate. But that's equally true with US frontier models. You can optimize your work by choosing between Opus, Fable, Sonnet, Sol, Luna and Terra for each task. Some people seem to prefer to let Opus code and Sol review, for example. And then there is the whole debate whether $current_version is actually better (Some people stay on Claude 4.8 because they dislike how 5.0 is sometimes doing stupid things, just as many opted out of dynamic reasoning when they still could)