Live data from Hacker News

DeepSeek V4 Pro 0813

openrouter.ai

51–60 of 493 posts

Re: DeepSeek V4 Pro 0813

#51
post #2

https://api-docs.deepseek.com/quick_start/pricing/ Competitive with opus 4.8 but weaker than sol or fable. About 20x cheaper.

Per token. You need to look at pricing per task.

... which still comes out cheaper, since DeepSeek caches so much more.

I keep track of my token consumption even on subscription plans and my equiv. cost for my 5.6-Sol usage is around $4000-$8000 a month.

Re: DeepSeek V4 Pro 0813

#52
post #7

Benchmarks: | Benchmark | DS-V4-Pro | DS-V4-Flash | DS-V4-Pro | DS-V4-Flash | GLM-5.2 | Kimi-K3 | Opus-4.8 | Fable 5 | | | 0813 | 0731 | Preview | Preview | | | | (w/ fallback) | |--------------------------|-----------|-------------|-----------|-------------|-----------|-----------|-----------|---------------| | HLE (wo/w tools) | 42.7/60.0 | 37.8/51.5 | 37.7/48.2 | 34.8/45.1 | 40.5/54.7 | 43.5/56.0 | 49.8/57.9 | 53.…

The timing looks like they are trying to take the wind out of Qwen's sails by releasing this on the same day that Qwen released the weights of Qwen3.8-max. Or maybe it's coincidence... For comparison I looked at Qwen's claimed benchmarks for Qwen3.8-max ( https://qwen.ai/blog?id=qwen3.8 ). Assuming each published set of benchmarks is believable, it looks like v4 Pro 0813 is better on average but overall performance i…

I mean at the rate of model releases happening, I think a lot of these will collide more often than expected!

Re: DeepSeek V4 Pro 0813

#53
post #42
post #36

Earlier quoted context omitted.

So it's a Fable class LLM? DSV4Pro vs Fable5 HLE w tools 60.0 vs 63.0 Terminal Bench 2.1 87.9 vs 88.0 Cybergym 83.3 vs 83.1 DeepSWE 62.7 vs 70.0 Toolathlon-Verified 74.1 vs 77.9 AutomationBench (Public) 31.8 vs 29.1 DSBench-FullStack 71.1 vs 77.2 DSBench-Hard 67.2 vs 68.3

Fable's guardrails would never let it do something like Cybergym so at least for that one it's measuring Opus 5

We have a first-party figure from the system card [1]:

> Mythos 5 reproduced 83.8% of targeted vulnerabilities on a single try, and produced at least one crash in 99.4% of tasks. This is comparable to Claude Mythos Preview, which reproduced 83.1% of targeted vulnerabilities and produced a crash in 97.1% of tasks. By contrast, Claude Opus 4.8 achieved a score of 78.1% (95.7% any crash).

So their quoted figure exactly matches the figure for Mythos Preview, although they don't state the provenance. It could also quite possibly be an independent measurement of Opus 5.

[1]: https://www-cdn.anthropic.com/57a52ea7d8f0e54e8a542e90826608...

Re: DeepSeek V4 Pro 0813

#54

Earlier quoted context omitted.

So not worth it over flash? Even at ~7x the size it isn't worth the price hike. Flash may be a monster of a model due to all the RL it received from free usage everywhere.

I haven't tried DeepSeek V4 Pro 0813 yet. Recent experience tells me that larger models are worth it in non-obvious ways. MiMo-V2.5-Pro solved problems that DeepSeek V4 Flash 0731 couldn't solve for me: for example, adding a live counter for elided reasoning lines to a terminal-based coding harness. You wouldn't be able to tell from the scores on their respective Artifical Analysis page ( https://artificialanalysis.a…

Interesting - I've been dropping into MiMo-V2.5-Pro-UltraSpeed whenever Flash seems to be "stuck" and it usually figures it out. I use UltraSpeed just because I'm so frustrated by then that I'm impatient.

I still find 5.6-Sol can solve some things neither of those can, but it's so slow (and it's so hard to trace / debug the reasoning) that I just let it run overnight.

Re: DeepSeek V4 Pro 0813

#55

Earlier quoted context omitted.

The timing looks like they are trying to take the wind out of Qwen's sails by releasing this on the same day that Qwen released the weights of Qwen3.8-max. Or maybe it's coincidence... For comparison I looked at Qwen's claimed benchmarks for Qwen3.8-max ( https://qwen.ai/blog?id=qwen3.8 ). Assuming each published set of benchmarks is believable, it looks like v4 Pro 0813 is better on average but overall performance i…

By that standard, the release of Grok 4.6 was also timed on the same day. Given how I think DeepSeek operates... I think they just release it when they feel it's ready, and don't even seem that concerned with what other people are doing.

Their leaks would confirm this sort of attitude. They're not trying to become the top player or anything like that - just working to play their part in pushing LLM tech forward and going from there. It was quite refreshing from the 'here's how we're going to dominate the world' nonsense. It's undoubtedly the same attitude that just lets them shrug and cancel the fund raising round after the leaks came from said funding round.

Re: DeepSeek V4 Pro 0813

#56
post #8

Earlier quoted context omitted.

Around 5 percentage points better. (E.g., 87% instead of 82%)

So not worth it over flash? Even at ~7x the size it isn't worth the price hike. Flash may be a monster of a model due to all the RL it received from free usage everywhere.

Opus 5 medium to Opus 5 max is only 3 points, if that puts it in context

Re: DeepSeek V4 Pro 0813

#57
post #33

Earlier quoted context omitted.

https://api-docs.deepseek.com/quick_start/pricing/ edit: there are banner announcements saying v4 flash pricing will increase first then overall by an undetermined amount

i dont see any price increase there... what am i missing?

Right below the pricing it is stated that they plan to increase the prices in the near future.

Re: DeepSeek V4 Pro 0813

#58
post #33

Earlier quoted context omitted.

https://api-docs.deepseek.com/quick_start/pricing/ edit: there are banner announcements saying v4 flash pricing will increase first then overall by an undetermined amount

i dont see any price increase there... what am i missing?

It's a big confusion, some[0] say an email was sent about significant price increase, personal I haven't seen anything official

[0] https://finance.yahoo.com/technology/ai/articles/deepseek-pl...

Re: DeepSeek V4 Pro 0813

#60
post #49

Earlier quoted context omitted.

The timing looks like they are trying to take the wind out of Qwen's sails by releasing this on the same day that Qwen released the weights of Qwen3.8-max. Or maybe it's coincidence... For comparison I looked at Qwen's claimed benchmarks for Qwen3.8-max ( https://qwen.ai/blog?id=qwen3.8 ). Assuming each published set of benchmarks is believable, it looks like v4 Pro 0813 is better on average but overall performance i…

Official pricing only kinda matters for an open weight model, no?

It still matters as a point of comparison until other providers come online. If the consensus price from other providers is much different that can be compared then. But for now we have $0.435 / $0.87 for v4 Pro 0813 (with increase announced but we don't know the new pricing), and $2 / $6 for Qwen3.8-max. So until we get other data points that is what we have to look at.
Post reply on HN