Live data from Hacker News

DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis

artificialanalysis.ai

181–190 of 342 posts

Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis

#182

> For the Code Agent tasks among the public benchmarks above, DeepSeek-V4-Flash-0731 is evaluated with the minimal mode of DeepSeek Harness (to be released) as the agent framework So, are they planning to announce an optimized coding agent harness as well ? DSv4 flash is a fantastic model, and my daily driver. With reasonix or pi, I can code all day long and pay a few pennies for it. No token anxiety. Whereas the sam…

If you're not paying for the model through Open router, then which provider are you buying it from? I thought that if that provider was listed on open router it would be the same price as going directly to the provider? Are you buying from a provider that isn't listed on open router already?

Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis

#183
post #85

So GLM 5.2/Gemini 3.6 level intelligence for $0.28/m output. And their updated Pro model coming soon.... Plus a size you can genuinely run at home: Unsloth lossless Q8 at 162GB.

I would like to see your "home"

Two (linked) DGX Sparks would do it I guess. Though probably slowly (I'd guess 15-20 tok/sec for decode, but higher for prefill). So ~$8-9k USD at current RAM prices, substantially less if they ever (sigh) drop. Electricity use would actually be relatively modest.

But it makes little to no sense as long as API prices are what they are. Except for maybe privacy reasons.

Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis

#187
post #59

Earlier quoted context omitted.

Because reddit unironically has better decorum around usage of their upvote/downvote system than HN does. People on HN downvote objectively correct information because they don't like it 24/7. There's a reason the creator of Zig left and gave the computer version of a middle finger on the way out to HN!

fwiw pg said early on that downvoting for disagreement is perfectly fine: https://news.ycombinator.com/item?id=117171 commenting about voting is also something the HN guidelines warns against: > Please don't comment about the voting on comments. It never does any good, and it makes boring reading. https://news.ycombinator.com/newsguidelines.html

Reminds me of telling employees to not discuss their wages. If voting here is such a toxic experience, maybe HN needs to wake up and smell the garbage.

Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis

#188

Earlier quoted context omitted.

Philosophically, I don’t believe we should outsource the accuracy of scripture to any single entity (let alone a for profit secular one). So it’s less about model choice and more about governance of scripture. I will check out the link you sent for sure!

I was under the impression you were getting BOOK.CHAPTER.VERSE references from the model and then sourcing them from a ground truth db. Anyways, impressive app! We haven't tackled such an ambitious project just for it being daunting.

It can still generate a wrong reference and select the wrong verse.

Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis

#189

> For the Code Agent tasks among the public benchmarks above, DeepSeek-V4-Flash-0731 is evaluated with the minimal mode of DeepSeek Harness (to be released) as the agent framework So, are they planning to announce an optimized coding agent harness as well ? DSv4 flash is a fantastic model, and my daily driver. With reasonix or pi, I can code all day long and pay a few pennies for it. No token anxiety. Whereas the sam…

They announced it already, read the tech report. "For the Code Agent tasks among the public benchmarks above, DeepSeek-V4-Flash-0731 is evaluated with the minimal mode of DeepSeek Harness (to be released) as the agent framework , using the max reasoning effort level with temperature = 1.0, top_p = 0.95."

> They announced it already

I meant something I could download and run.

Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis

#190

Earlier quoted context omitted.

Good luck when the last non-Chinese frontier labs will have closed and the CCP will ask to stop sharing models open source.

My conspiracy theory is that this is the new space race, and the CCP encourages this to show the world what Chinese engineers are capable of, and tank the Anthropic/OpenAI valuation bubble as a desirable side effect.

Half of the engineers at OAI and Anthropic are Asian, I don't think China does all these for signaling.
Post reply on HN