Live data from Hacker News

DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis

artificialanalysis.ai

81–90 of 342 posts

Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis

#81

New Deepseek models are like Christmas for me. Really big fan of low cost API models, noone does it better than DS. Until VRAM price is low enough to run models locally, this is the way to go. The subsidized subscription model won't last, API pricing "feels" closer to a true sustainable business model.

Indeed. My fellow software engineers keep complaining about using up all their Claude tokens within an hour... Whilst I'll be rocking DS flash for the entire day. Sure it gets a few things wrong here and there, but that's when you pull out the Claude models or whatever for those tricky tasks.

Same. And I have come to use OMP (oh my pi) agent /advisor mode to put a 2nd model on the case (also mid-size one), reading everything. It can not block anything or change anything - just inserts comments in the text stream with 1 turn delay. Good portion of the time it's quiet. I'd say 1/2 of the time it's got something to say. About 2/3-rd of that the 'advice' is insubstantial or about something not-quite wrong. The good thing is the main model is confident - checks and then it stands its ground. Have not noticed it turning a right into a wrong b/c of advisor false alarm. And in 1/3-rd of the advice, it's a genuine defect teh advisor noticed, the main model works out a fix. This is my approximate feeling just observing the process, have not got collected the data. Afaik only OMP has advisor mode. Agent pi has plugin pi-omplike-advisor. For agent Hermes I had them code me an /advisor plugin (for now -0.1 old v0.18.x; yet to upgrade it to latest).

Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis

#82
post #20

If deepseek v4 flash is beating DeepSeek V4 Pro, can we expect new V4 Pro which is on par with Opus 5 in couple weeks (even better if it beats Opus)?

I’ve been using v4 flash for an app I’m building [1] and it’s amazing how cost effective and good it is coming from having always used gpt, opus and sonnet models. It’s so cost effective I can offer a generous free tier since my goal isn’t to make money with it. [1] https://trysojourn.app

Is the difference between this and a frontier model that the scripture is guaranteed to be real?

I'm on a team that develops a Bible study app, and we're all relatively content with how the basic models converse regarding scripture. Even as far back as GPT-4 was excellent. They occasionally have minor hallucinations (a dealbreaker for a production app), but they do an excellent job with theology and Bible scholarship, given reasonable guardrails.

I'll admit I'm coming from the perspective of "should we be implementing this?" It seems, on the surface, that a strong embedding-based verse retrieval covers the bases at a microfraction of the cost.

If you're interested, check out the development server where we're working on this. You navigate to the search (magnifying glass) and then hit "Meaning". Sorry for the confusing route; we're still deciding on back-end details and haven't focused on the front yet.

[1] https://ai.stepbible.org

Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis

#84
post #36

Earlier quoted context omitted.

By making companies using them "toxic" to touch. For example: no government contract to any company who uses even one vendor in it's entire chain of dependencies, who uses such open models. They can extend this further by laying more conditions, such as: any company dealing in this-this field can only use models "officially" approved as "safe". Rest you can guess how easy it would be to get that "safe" rating for suc…

So now the US companies will be stuck on expensive models while the rest of the world can do things much more cost effective. I'm not sure the outcome would be beneficial for the US as a whole here. But perhaps that is not their priority.

There are many proxy-ways to bypass such restrictions. Of course, the costs will be higher. Providers will spun-up.

Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis

#86

The weights were just released a few minutes ago: https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731

Can't wait for the DwarfStar quants - I have been using DeepSeek v4 flash (preview) as my main coding agent for months now (running on my 128gb mbp) - it seems this model outperforms GLM 5.2 on nearly every metric. Thanks for sharing the news, I was refreshing huggingface but gave up thinking it likely would take some more time.

What kind of tps are you getting?

Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis

#87
post #49

I’m wondering if they did anything to address the DSML tool calls leaking. Has been an issue with both Flash and Pro so far.

On OpenRouter it seemed to only affect a couple of specific providers. Ignoring them has made it s non issue for me.

Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis

#89

Earlier quoted context omitted.

Genocide in Gaza...

I find that ChatGPT isn't censoring, but it is being pretty weaselly. If you ask it "is there genocide in gaza". It will say no but also say that a lot of organizations classify it as such. It will then say "it's highly disputed". If you poke it just a few times, however, you get to the point where it will eventually say (paraphrasing) that basically only Israel, the US state department, and the ICJ say it's not a ge…

One side of a conflict being a minority does not make it less nuanced.

The majority of the world is religious - doesn’t mean the debate on religion isn’t a complex question.

The majority of the world approved of slavery historically.

The majority of countries have ethnically cleansed their Jews, many of them in living memory.

When interrogated you will find that the only ones asserting the war in Gaza is a genocide are people who were anti-Israel anyway.

Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis

#90

Earlier quoted context omitted.

Indeed. My fellow software engineers keep complaining about using up all their Claude tokens within an hour... Whilst I'll be rocking DS flash for the entire day. Sure it gets a few things wrong here and there, but that's when you pull out the Claude models or whatever for those tricky tasks.

what plan are your 'fellow software engineers' using? I have a hard time even using up the Fable part of my allowance in a week of coding.

I'm not sure to be fair, but they do have constant "token anxiety", which I simply don't have anymore since using v4 flash.
Post reply on HN