Live data from Hacker News

DeepSeek-V4-Flash Update

api-docs.deepseek.com

1–10 of 362 posts

Re: DeepSeek-V4-Flash Update

#2
DeepSeek V4 Flash (Preview → 2026-07-31)

• Terminal Bench: 56.9 → 82.7 (+25.8)

• Toolathlon: 51.8 → 70.3 (+18.5)

Compared to GPT-5.6 Terra:

• Terminal Bench: Flash 82.7 vs Terra 78.4

• Toolathlon: Flash 70.3 vs Terra 53.1

• DeepSWE: Flash 54.4 vs Terra 69.6

• Agents' Last Exam: Flash 25.2 vs Terra 50.4

Trading blows with Terra, which is pretty interesting. No clear winner on these benchmarks, and wildy differeing scores. Very interesting!

Re: DeepSeek-V4-Flash Update

#3
This is more exciting than k3, IMO. Dsv4 models are extremely cheap to serve. Improving their capabilities has lots of downstream effects, as it becomes "good enough" for more and more tasks.

DS was serving the pro version at extremely low prices for a long time, and they've had integrations with opencode & other providers, so they likely gathered a lot of data from real developers doing real tasks (on openrouter they were labeled as such). Now they can use those live scenarios to further post-train their models and improve them further.

Can't wait to see if distilling k3 into dsv4 brings additional improvements. Anyway, having fast cheap models getting better is great for the community. Especially since these don't "go away" on a provider's whim. Whatever capabilities they get, can be used "forever" going forward. And, at least flash can be ran "at home" with <10k in hardware, which isn't really possible / feasible with glm/k3 larger models.

Re: DeepSeek-V4-Flash Update

#4
In case people want to run it, it's DeepSeek-V4-Flash-284B-A13B. So it should just barely run on a single B300, and it's small enough that it'll barely run on an M5 Max too.

Re: DeepSeek-V4-Flash Update

#6
post #4

In case people want to run it, it's DeepSeek-V4-Flash-284B-A13B. So it should just barely run on a single B300, and it's small enough that it'll barely run on an M5 Max too.

It runs really well on 2 DGX Sparks - 60t/s

Re: DeepSeek-V4-Flash Update

#7
post #2

DeepSeek V4 Flash (Preview → 2026-07-31) • Terminal Bench: 56.9 → 82.7 (+25.8) • Toolathlon: 51.8 → 70.3 (+18.5) Compared to GPT-5.6 Terra: • Terminal Bench: Flash 82.7 vs Terra 78.4 • Toolathlon: Flash 70.3 vs Terra 53.1 • DeepSWE: Flash 54.4 vs Terra 69.6 • Agents' Last Exam: Flash 25.2 vs Terra 50.4 Trading blows with Terra, which is pretty interesting. No clear winner on these benchmarks, and wildy differeing sco…

Since they did this with their own harness I m not sure it's apples to apples comparison.

Re: DeepSeek-V4-Flash Update

#8
post #7
post #2

DeepSeek V4 Flash (Preview → 2026-07-31) • Terminal Bench: 56.9 → 82.7 (+25.8) • Toolathlon: 51.8 → 70.3 (+18.5) Compared to GPT-5.6 Terra: • Terminal Bench: Flash 82.7 vs Terra 78.4 • Toolathlon: Flash 70.3 vs Terra 53.1 • DeepSWE: Flash 54.4 vs Terra 69.6 • Agents' Last Exam: Flash 25.2 vs Terra 50.4 Trading blows with Terra, which is pretty interesting. No clear winner on these benchmarks, and wildy differeing sco…

Since they did this with their own harness I m not sure it's apples to apples comparison.

> not sure it's apples to apples comparison.

They're literally comparing the previous version of the same model with the new one. It's based on the same architecture, same pre-trained model, just different post-training. It doesn't get more apples to apples than this.

Re: DeepSeek-V4-Flash Update

#9
post #7
post #2

DeepSeek V4 Flash (Preview → 2026-07-31) • Terminal Bench: 56.9 → 82.7 (+25.8) • Toolathlon: 51.8 → 70.3 (+18.5) Compared to GPT-5.6 Terra: • Terminal Bench: Flash 82.7 vs Terra 78.4 • Toolathlon: Flash 70.3 vs Terra 53.1 • DeepSWE: Flash 54.4 vs Terra 69.6 • Agents' Last Exam: Flash 25.2 vs Terra 50.4 Trading blows with Terra, which is pretty interesting. No clear winner on these benchmarks, and wildy differeing sco…

Since they did this with their own harness I m not sure it's apples to apples comparison.

It will be fair if they release the harness though. I think now the future will be paired model-harness releases, not just weight dumps.

The performance changes are so big with the right harness that is makes sense to engineer the harness and fine-tune the model to one another from the start.

Re: DeepSeek-V4-Flash Update

#10
post #7

Earlier quoted context omitted.

Since they did this with their own harness I m not sure it's apples to apples comparison.

> not sure it's apples to apples comparison. They're literally comparing the previous version of the same model with the new one. It's based on the same architecture, same pre-trained model, just different post-training. It doesn't get more apples to apples than this.

I think the commenter means the Flash vs Terra benchmarks.
Post reply on HN