Live data from Hacker News

DeepSeek V4 Pro 0813

openrouter.ai

91–100 of 493 posts

Re: DeepSeek V4 Pro 0813

#91
post #7

Benchmarks: | Benchmark | DS-V4-Pro | DS-V4-Flash | DS-V4-Pro | DS-V4-Flash | GLM-5.2 | Kimi-K3 | Opus-4.8 | Fable 5 | | | 0813 | 0731 | Preview | Preview | | | | (w/ fallback) | |--------------------------|-----------|-------------|-----------|-------------|-----------|-----------|-----------|---------------| | HLE (wo/w tools) | 42.7/60.0 | 37.8/51.5 | 37.7/48.2 | 34.8/45.1 | 40.5/54.7 | 43.5/56.0 | 49.8/57.9 | 53.…

Geometric mean of all these benchmarks : * GPT-5.6 Sol: 65.5 * Fable 5 (w/ fallback): 64.5 * Opus 5: 64.0 * DS-V4-Pro 0813: 62.5 * Kimi-K3: 62.3 * DS-V4-Flash 0731: 55.8 * GLM-5.2: 47.3

Maybe it's me but I don't see how DS Flash is better than GLM at all, much less by a huge gap. I'd probably protest less against Fable and Opus being put at the same level than many would, but there's no denying the two models are a very different experience from each other. I guess where I'm going is no one should pick a model by the benchmarks.

Re: DeepSeek V4 Pro 0813

#92

Earlier quoted context omitted.

Their leaks would confirm this sort of attitude. They're not trying to become the top player or anything like that - just working to play their part in pushing LLM tech forward and going from there. It was quite refreshing from the 'here's how we're going to dominate the world' nonsense. It's undoubtedly the same attitude that just lets them shrug and cancel the fund raising round after the leaks came from said fundi…

The founder of DS's stated goal is to get to AGI. He thinks this is the path to get there. Kind of interesting, when compared to the hubris from American frontier labs.

> Kind of interesting, when compared to the hubris from American frontier labs.

One Man’s “hubris” is another man’s “marketing campaign.”

Drama sells.

Re: DeepSeek V4 Pro 0813

#93
post #7

Benchmarks: | Benchmark | DS-V4-Pro | DS-V4-Flash | DS-V4-Pro | DS-V4-Flash | GLM-5.2 | Kimi-K3 | Opus-4.8 | Fable 5 | | | 0813 | 0731 | Preview | Preview | | | | (w/ fallback) | |--------------------------|-----------|-------------|-----------|-------------|-----------|-----------|-----------|---------------| | HLE (wo/w tools) | 42.7/60.0 | 37.8/51.5 | 37.7/48.2 | 34.8/45.1 | 40.5/54.7 | 43.5/56.0 | 49.8/57.9 | 53.…

In classic reddit fashion the post you linked to is now deleted

Re: DeepSeek V4 Pro 0813

#94
post #68

Tested both DS v4 pro 0813 and Grok 4.6 (all from openrouter) on Codex cli. Worked on a same new feature development on my project. Deepseek 4 pro: Worked for 12m 02s - cost $0.12 - has bug. Grok 4.6: Worked for 3m 18s - cost $ 1.41 - no bug.

[flagged]

no he and his stuffs are now considered transparent, no pun intended. I think he deserves it since his minions were persistent with usage of "this ___ has hateful bias against ___" canned response.

Re: DeepSeek V4 Pro 0813

#95

Earlier quoted context omitted.

Repeat the test like 5 times for each model and see the results.

+1, a single test means little.

I don't think so. I specifically kept this PR to test model capabilities, and I've already tested a bunch of models. Current test results show that the more advanced the model is, the easier it passes. For example, GPT-5.5 Medium fails the test(has bug), but High passed.

Re: DeepSeek V4 Pro 0813

#96
post #50

I've been using the last Deepseek Flash update for a week and I'm amazed. It was a capable model for easy tasks but now it looks like it can do some heavy development for peanuts. I can't wait to try this new one.

IME I can't trust it to write it's own plans from a spec, but if I give it a detailed execution plan written by Opus, it's fast and cheap (if chatty) in executing it.

Re: DeepSeek V4 Pro 0813

#98

Tested both DS v4 pro 0813 and Grok 4.6 (all from openrouter) on Codex cli. Worked on a same new feature development on my project. Deepseek 4 pro: Worked for 12m 02s - cost $0.12 - has bug. Grok 4.6: Worked for 3m 18s - cost $ 1.41 - no bug.

I thought it was impossible to downvote posts?

Re: DeepSeek V4 Pro 0813

#99
post #93
post #7

Benchmarks: | Benchmark | DS-V4-Pro | DS-V4-Flash | DS-V4-Pro | DS-V4-Flash | GLM-5.2 | Kimi-K3 | Opus-4.8 | Fable 5 | | | 0813 | 0731 | Preview | Preview | | | | (w/ fallback) | |--------------------------|-----------|-------------|-----------|-------------|-----------|-----------|-----------|---------------| | HLE (wo/w tools) | 42.7/60.0 | 37.8/51.5 | 37.7/48.2 | 34.8/45.1 | 40.5/54.7 | 43.5/56.0 | 49.8/57.9 | 53.…

In classic reddit fashion the post you linked to is now deleted

To be fair… I don’t know who still needs to figure out that AI benchmarks are almost all entirely fucking trash, but the great number would surely surprise me.

Re: DeepSeek V4 Pro 0813

#100

Tested both DS v4 pro 0813 and Grok 4.6 (all from openrouter) on Codex cli. Worked on a same new feature development on my project. Deepseek 4 pro: Worked for 12m 02s - cost $0.12 - has bug. Grok 4.6: Worked for 3m 18s - cost $ 1.41 - no bug.

[deleted]
Post reply on HN