Live data from Hacker News

DeepSeek-V4-Flash Update

api-docs.deepseek.com

191–200 of 362 posts

Re: DeepSeek-V4-Flash Update

#191

Earlier quoted context omitted.

An experienced software engineer (read his profile) praises the value he's found in Deepseek, and gives some real data showing how affordable that value is. Then you, dakolli - out of generosity and minute-to-minute devotion to enlightenment - sacrifice time from your busy day to sit down (though perhaps that's been painful lately?) or stand up with your phone - and offer a profound, deeply thought-out counterpoint i…

[flagged]

I agree a bio is a bad proxy, and I'm not saying I'm amazing but I'm not THAT incompetent. Most of my GitHub is pre-LLM era; github.com/lionkor

Edit: And like a lot of tools, HOW you use them is just about as important as the quality of the tool itself. Of course the tools can produce massive amounts of bad quality slop, they can also produce fast, focused edits that make sense.

Re: DeepSeek-V4-Flash Update

#192
post #149

Earlier quoted context omitted.

Qwen has a pinned tweet stating that 3.8 will be released as open weights soon. I guess it remains to be seen, though, if they’ll do the smaller model sizes or only the big 2.4T one.

Most interesting part IMO is whatever Qwen open weight will include image / video capabilities since its where they are strongest.

Also base models versions are very useful for researchers, since they allow for a cold start to post training.

Re: DeepSeek-V4-Flash Update

#194

What was preventing them from calling it v4.1-Flash to distinguish it better?

Sounds like it'll replace v4-flash, v4.1 would be nice to keep both available. On the other hand, it's nice to just get an improvement on anything that asks for "deepseek-v4-flash" without having to change the model string.

Re: DeepSeek-V4-Flash Update

#196
post #18

Woah, a 200B model competing with GLM-5.2 and getting close to Opus 4.8. Quite impressive. If those numbers translate well to its general capabilities, with the great caching DeepSeek has, I feel like this model will get tons of usage.

MiniMax's another lab that's known for relatively smaller models (their latest, M3 is 295b) that punch way above its weight.

Re: DeepSeek-V4-Flash Update

#197

Earlier quoted context omitted.

An experienced software engineer (read his profile) praises the value he's found in Deepseek, and gives some real data showing how affordable that value is. Then you, dakolli - out of generosity and minute-to-minute devotion to enlightenment - sacrifice time from your busy day to sit down (though perhaps that's been painful lately?) or stand up with your phone - and offer a profound, deeply thought-out counterpoint i…

[flagged]

> I really can't begin to describe how stupid I think you are for thinking that having that statement in your bio makes you a qualified engineer.

Well, I guess having "Two PhDs, Three masters degrees. Expert in everything.." in your bio makes you very smart!

Re: DeepSeek-V4-Flash Update

#198
post #91

Deepseek and moonshot are the only two providers I consent to training for.

Xi going to shut down open-weighting of them in a matter of months. No one seriously doubts this. They will be too powerful and they will be gone.

Reuters was reporting that rumor. And then Xi made a public appearance at a conference in Shanghai where he said the opposite of that rumor.

Re: DeepSeek-V4-Flash Update

#199
post #98

Earlier quoted context omitted.

On DeepSWE Deepseek is 54.4% and Luna is 67%

But on Terminal bench, its * DS4 Flash: 82.7 * GPT 5.6 Luna: 75.7 For reference, that puts it on the third spot behind GPT 5.5 and Fable 5. For some reason GPT 5.6 Sol is not showing in the leaderboard. If it did, then DS4 Flash was number four. The thing is, even if Luna is better in DeepSWE and has the 80% discount. DeepSeek is still cheaper. -------------- DeepSeek V4 | Flash GPT-5.6 Luna (New) -------------- Inpu…

Agreed. The only thing missing is multimodal capability

Re: DeepSeek-V4-Flash Update

#200

Earlier quoted context omitted.

Gpt 5.6 Luna cache read is $0.02 per mtok V4 flash cache read is $0.0028 per mtok That's not "a bit cheaper", just saying

That's a good point. Yeah their caching input is insane.

DeepSeek v4 API rates are only matched by Xiaomi's MiMo v2.5 series. Quite similar in capabilities too (at least before today's update).
Post reply on HN