Live data from Hacker News

DeepSeek V4 Pro 0813

openrouter.ai

181–190 of 493 posts

Re: DeepSeek V4 Pro 0813

#181
post #67

Earlier quoted context omitted.

Their leaks would confirm this sort of attitude. They're not trying to become the top player or anything like that - just working to play their part in pushing LLM tech forward and going from there. It was quite refreshing from the 'here's how we're going to dominate the world' nonsense. It's undoubtedly the same attitude that just lets them shrug and cancel the fund raising round after the leaks came from said fundi…

Benefits of having a well performing hedge fund funding DeepSeek. IIRC, Demis attempted to start a fund inside DeepMind but it was killed off. In an alternative world where he manages to pull that off, perhaps DeepMind would still be independent with Demis at the helm.

that if fund would be profitable.

Re: DeepSeek V4 Pro 0813

#182

Earlier quoted context omitted.

Geometric mean of all these benchmarks : * GPT-5.6 Sol: 65.5 * Fable 5 (w/ fallback): 64.5 * Opus 5: 64.0 * DS-V4-Pro 0813: 62.5 * Kimi-K3: 62.3 * DS-V4-Flash 0731: 55.8 * GLM-5.2: 47.3

Maybe it's me but I don't see how DS Flash is better than GLM at all, much less by a huge gap. I'd probably protest less against Fable and Opus being put at the same level than many would, but there's no denying the two models are a very different experience from each other. I guess where I'm going is no one should pick a model by the benchmarks.

It isn't. I run both at home. GLM5.2 Q4 crushes DSv4Flash0731 Q8. I reach for DS for speed and for medium effort level work. If I care about quality I'll reach for GLM5.2 Looking at this release, I'm comparing it to GLM5.2 and it seems to beat GLM5.2, only time/experience will show. If true, then I'm happy. It's much easier to run than Qwen3.8/KimiK3

Re: DeepSeek V4 Pro 0813

#183

Earlier quoted context omitted.

Per token. You need to look at pricing per task.

... which still comes out cheaper, since DeepSeek caches so much more. I keep track of my token consumption even on subscription plans and my equiv. cost for my 5.6-Sol usage is around $4000-$8000 a month.

How much do you pay for the subscription?

Re: DeepSeek V4 Pro 0813

#184
post #153

Earlier quoted context omitted.

I think I saw a better overall composition out of Flash 0731 Effort on this one?

Default effort for OpenRouter. I'll try a grid of efforts... Wow, the low, medium, and high pelicans came out in surprisingly different styles: https://tools.simonwillison.net/markdown-svg-renderer#url=ht...

exciting. it's almost like 3 models in one. that variety would matter when trying to solve a creative problem.

Re: DeepSeek V4 Pro 0813

#185
post #127
post #121

Earlier quoted context omitted.

wait for Deepseek Harness (yes it is the official name) release then try again. for your kind of task, harness tools matter.

If the model cannot figure out simple and ubiquitous tools, how is it supposed to figure out complex problems? All of the good models basically work with any harness, including giving them a single "shell command" tool. They can just figure things out.

the single shell command is the terminal bench.

Re: DeepSeek V4 Pro 0813

#186

Earlier quoted context omitted.

[flagged]

There are still sane people here, we just don't talk about "rocket man" because it enrages the particular subgroup on display here and usually goes no where actually productive (and has like a 50% of getting flagged to death anyways). I'm not pro Elon by any means, but the standard HN profile of him is pretty bat shit crazy.

He is honestly batshit insane

Re: DeepSeek V4 Pro 0813

#188

Just tested through openrouter.. gave exactly same task.. the task was to scan existing repo, and generate a single docker-compose file to deploy behind a caddy server, where certain port ranges are already used, the service demands widlcard certificates to be provisioned from outside, and postgre needs to be built-in one... Tested this model, and gpt-5.6-terra-high. Results: this one had few issues. terra: none. The…

I use deepseek flash to do exactly this. Git repo (which I usually have it build from scratch) -> build docker image -> deploy to server with komodo/caddy-docker proxy.

Works great, regularly one shot applications. I often make changes to the application after its deployed (to be fair, my prompts are usually quite laxidasical, just 'build x, use /deploy-to-komodo) but the deployment works great.

I did make a skill, but if your doing anything repeatedly you should as well.

Opencode, but any harness I'd think would work similar.

Re: DeepSeek V4 Pro 0813

#189
post #123

Earlier quoted context omitted.

What harness are you using? DS V4 is harness sensitive.

Are we testing the model or the harness? If benchmarks show certain numbers it should perform as such without it

You can't even run a benchmark without a basic harness, of course the harness has some effect.

Re: DeepSeek V4 Pro 0813

#190

Worse than Luna but more expensive than Luna. Sticking with Luna without sending my data to Deepseek (China)

Unless you're Chinese, why would you care if they see your data? As an American, I'd much rather have my data kept outside the country than here where companies and the government have a lot more leverage over me.

I do think this question comes up a lot-- I can understand why.

For some well-explained reasons, check out https://darioamodei.com/essay/the-adolescence-of-technology and search for "CCP".

Post reply on HN