Live data from Hacker News

DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis

artificialanalysis.ai

161–170 of 342 posts

Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis

#161

The weights were just released a few minutes ago: https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731

Can't wait for the DwarfStar quants - I have been using DeepSeek v4 flash (preview) as my main coding agent for months now (running on my 128gb mbp) - it seems this model outperforms GLM 5.2 on nearly every metric. Thanks for sharing the news, I was refreshing huggingface but gave up thinking it likely would take some more time.

[deleted]

Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis

#164
post #12

Does it already know the answer to what happen at Tiananmen Square? Or still avoiding it?

How often are you asking an LLM about this? You know you can just google it, right? I have to admit it rarely comes up in the coding tasks I usually give to LLMs.

Surely you realize he or she is not literally looking for the answer to that but rather pointing out the censorship.

Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis

#165
> For the Code Agent tasks among the public benchmarks above, DeepSeek-V4-Flash-0731 is evaluated with the minimal mode of DeepSeek Harness (to be released) as the agent framework

So, are they planning to announce an optimized coding agent harness as well ? DSv4 flash is a fantastic model, and my daily driver. With reasonix or pi, I can code all day long and pay a few pennies for it. No token anxiety. Whereas the same model with fireworks/openrouter, with zdr thrown in, token costs ratchet up with no explanation. Likely that the model is subsidized for gathering usage data. I am waiting for the day I can run this locally.

Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis

#166

Earlier quoted context omitted.

The thing is, you don’t need to actually block usage to make something illegal. You make it so toxic that company wants to be seen publicly using open models

Companies would secretly use self-hosted models internally because it would give them a enormous cost advantage.

Sure, companies can indeed do illegal things

Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis

#167

Earlier quoted context omitted.

Indeed. My fellow software engineers keep complaining about using up all their Claude tokens within an hour... Whilst I'll be rocking DS flash for the entire day. Sure it gets a few things wrong here and there, but that's when you pull out the Claude models or whatever for those tricky tasks.

The problem is picking between models. I do not want to spend my time switching models and trying to decipher which model should be used for what. Maybe that's just a me problem that I need to figure out.

DeepSeek v4 is honestly good enough that I'm fine throwing it at everything in my hobby projects. I guess now I'll be switching from V4 Pro-Preview to V4 Flash. My only real complaint is that they can't do images, which limits their ability to autonomously debug some kinds of issues

Of course you can get more bang for your buck by being more deliberate. But that's equally true with US frontier models. You can optimize your work by choosing between Opus, Fable, Sonnet, Sol, Luna and Terra for each task. Some people seem to prefer to let Opus code and Sol review, for example. And then there is the whole debate whether $current_version is actually better (Some people stay on Claude 4.8 because they dislike how 5.0 is sometimes doing stupid things, just as many opted out of dynamic reasoning when they still could)

Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis

#170
I can't help but feel the timing coinciding with luna's price updates to be somewhat strategic. But without multi-modality it's slightly dead-in-the-water for my usecase: https://design.withfudge.com. I'm currently using Minimax-M3, but Luna ekes out abit futher on the intelligence, so i'll be switching to it very soon.
Post reply on HN