Live data from Hacker News

DeepSeek V4 Pro 0813

openrouter.ai

81–90 of 493 posts

Re: DeepSeek V4 Pro 0813

#81
post #6

I find it interesting how much adoption seems to be influenced by momentum. Some of these Chinese models are surprisingly capable, but developers often default to the models that are already established as the “industry standard

Most of my model usage comes from my work’s model selection (which is now down to just Claude models)

I’ll try out the latest models, but mainly stick with Claude only because I’m most used to its quirks and how to work around them. I imagine this is part of these hyperscalers playbook.

I will say though, I miss Sol model at work. It with Codex was amazing at first-shot understanding. Claude i need to scope out where to look otherwise a large portion of my token budget is eaten up

Re: DeepSeek V4 Pro 0813

#83

It appears that the only available endpoint (as of this writing) requires enabling "Allow paid endpoints that train on request data" in the OpenRouter privacy settings. I hope additional paid providers will become available that don't require training on data.

That is likely because Deepseek themselves is the only host.

In 24-48 hours there will be other options I presume

Re: DeepSeek V4 Pro 0813

#84

Currently burning money quickly on official deepseek api. They are also increasing pricing starting today. V4 Flash 0731 still feels like the most outstanding model of the past few months and probably to come.

Deepseek seems to have gotten too cheap. I have been using it for a long time and it's at a point now where my credits balance barely moves even at max setting.

Re: DeepSeek V4 Pro 0813

#85

Worse than Luna but more expensive than Luna. Sticking with Luna without sending my data to Deepseek (China)

Unless you're Chinese, why would you care if they see your data? As an American, I'd much rather have my data kept outside the country than here where companies and the government have a lot more leverage over me.

[deleted]

Re: DeepSeek V4 Pro 0813

#87
post #39
post #36

Earlier quoted context omitted.

So it's a Fable class LLM? DSV4Pro vs Fable5 HLE w tools 60.0 vs 63.0 Terminal Bench 2.1 87.9 vs 88.0 Cybergym 83.3 vs 83.1 DeepSWE 62.7 vs 70.0 Toolathlon-Verified 74.1 vs 77.9 AutomationBench (Public) 31.8 vs 29.1 DSBench-FullStack 71.1 vs 77.2 DSBench-Hard 67.2 vs 68.3

Fabble lol

[dead]

Re: DeepSeek V4 Pro 0813

#88

It appears that the only available endpoint (as of this writing) requires enabling "Allow paid endpoints that train on request data" in the OpenRouter privacy settings. I hope additional paid providers will become available that don't require training on data.

Their privacy policy doesn't forbid them from just straight up publishing your raw prompts as training data.

My threat model is that anything I POST to DeepSeek I treat as public to the web, as much as a public GitHub repo is.

Re: DeepSeek V4 Pro 0813

#89

What I care about is whether the model is capable of the tasks I give it at the lowest cost. Right now I'm using Kimi-K3/GLM-5.2/Minimax. Sonnet is great but I burn through the tokens too fast. Opus 5 set to max is amazing and more intelligent than all of us. .998 of the time I don't need that kind of intelligence. I just need the job done.

How do you define intelligence? I encounter that kind of sentiment all too often, and I have to assume we go by wildly different understanding of what that might entail.

If you read Opus 5's output, it is beyond the comprehension of virtually all engineers and developers. That is what I mean by intelligence. Math, science, and engineering are all contained in one model. We may be experts in one field. The model is an expert in everything that humans know.

Re: DeepSeek V4 Pro 0813

#90

Tested both DS v4 pro 0813 and Grok 4.6 (all from openrouter) on Codex cli. Worked on a same new feature development on my project. Deepseek 4 pro: Worked for 12m 02s - cost $0.12 - has bug. Grok 4.6: Worked for 3m 18s - cost $ 1.41 - no bug.

Repeat the test like 5 times for each model and see the results.

+1, a single test means little.
Post reply on HN