Benchmarks: | Benchmark | DS-V4-Pro | DS-V4-Flash | DS-V4-Pro | DS-V4-Flash | GLM-5.2 | Kimi-K3 | Opus-4.8 | Fable 5 | | | 0813 | 0731 | Preview | Preview | | | | (w/ fallback) | |--------------------------|-----------|-------------|-----------|-------------|-----------|-----------|-----------|---------------| | HLE (wo/w tools) | 42.7/60.0 | 37.8/51.5 | 37.7/48.2 | 34.8/45.1 | 40.5/54.7 | 43.5/56.0 | 49.8/57.9 | 53.…
So it's a Fable class LLM? DSV4Pro vs Fable5 HLE w tools 60.0 vs 63.0 Terminal Bench 2.1 87.9 vs 88.0 Cybergym 83.3 vs 83.1 DeepSWE 62.7 vs 70.0 Toolathlon-Verified 74.1 vs 77.9 AutomationBench (Public) 31.8 vs 29.1 DSBench-FullStack 71.1 vs 77.2 DSBench-Hard 67.2 vs 68.3
DeepSeek V4 Pro 0813
201–210 of 493 posts
Re: DeepSeek V4 Pro 0813
#202Earlier quoted context omitted.
What's the new pricing? The prices on OpenRouter still look the same.
nobody is saying. just "more". but openrouter says they don't expect the price to change other than through the deepseek api, other people hosting the same model will keep charging the same price.
Re: DeepSeek V4 Pro 0813
#203Earlier quoted context omitted.
If that wasn't impressive enough, it's actually ~60x cheaper if you take into account the typical cache-read/input/output split in agentic coding, and the deep discount for cache reads offered by DeepSeek. Opencode has some public data on the typical split [1]: For DeepSeek V4 Pro the typical split is 750 in, 290 out, 82k cached. Cost per request for V4 Pro: $0.000875 per request. Equivalent Opus cost (w/o taking int…
I created a simulation for coding harnesses based on my own pi sessions. When taking into account all factors, DS-v4-Pro is cheaper than gpt-5.6-luna due to caching. Look at the bill segments difference for cache read cost and uncached cost between deepseek and the other models. At this point is cheaper to use ds-v4-pro than the luna models from openai. ignore the numbers except the classic and keep in mind that clas…
Re: DeepSeek V4 Pro 0813
#204Earlier quoted context omitted.
... which still comes out cheaper, since DeepSeek caches so much more. I keep track of my token consumption even on subscription plans and my equiv. cost for my 5.6-Sol usage is around $4000-$8000 a month.
How much do you pay for the subscription?
Re: DeepSeek V4 Pro 0813
#205Earlier quoted context omitted.
I've found Pro to be a lot better per "task" than the recently released Flash for code reviews and things (via OpenRouter running in pi.dev). Flash makes a lot more initial mistakes, and then has to re-check stuff, and produces much more output compared to Pro. It often gets to the correct result eventually, but the output volume is often 5x more than for Pro, and the initial outputs are often wrong, with the first f…
yeah, I've definitely noticed one has to be quite precise to keep Flash on the straight-and-narrow
Re: DeepSeek V4 Pro 0813
#206Again, I will wait until there's a provider that doesn't train on prompts before I will benchmark.
Let’s just wait a bit for this one.
Re: DeepSeek V4 Pro 0813
#207Earlier quoted context omitted.
If you read Opus 5's output, it is beyond the comprehension of virtually all engineers and developers. That is what I mean by intelligence. Math, science, and engineering are all contained in one model. We may be experts in one field. The model is an expert in everything that humans know.
I'd have to ask for you to be more specific, otherwise, to take your answer at face value, it comes across as a contradiction. > [Opus 5's output] is beyond the comprehension of virtually all engineers and developers That would make it pretty bad? The key defining quality of good software, is clarity, and the ability to simplify a complex problem to the point of it seeming trivial. > Math, science, and engineering ar…
Re: DeepSeek V4 Pro 0813
#208Earlier quoted context omitted.
Are we testing the model or the harness? If benchmarks show certain numbers it should perform as such without it
You can't even run a benchmark without a basic harness, of course the harness has some effect.
Re: DeepSeek V4 Pro 0813
#209Earlier quoted context omitted.
Terra has not been able to do any of the technical tasks I've asked of it correctly. I'm surprised others get use out of it. Anything below Sol high tends to give me mostly unreliable results. I'm using codex as my main harness but maybe it performs better with a different one.
I bounce between Sol high/medium and Luna max. I don't know why you'd use anything between Luna max and Sol medium. Luna is so extremely cheap and cranked up to max it does anything I'd want Terra to do for a fraction of the cost. What is Terra for?
Re: DeepSeek V4 Pro 0813
#210Worse than Luna but more expensive than Luna. Sticking with Luna without sending my data to Deepseek (China)