Just tested through openrouter.. gave exactly same task.. the task was to scan existing repo, and generate a single docker-compose file to deploy behind a caddy server, where certain port ranges are already used, the service demands widlcard certificates to be provisioned from outside, and postgre needs to be built-in one... Tested this model, and gpt-5.6-terra-high. Results: this one had few issues. terra: none. The…
DeepSeek V4 Pro 0813
171–180 of 493 posts
Re: DeepSeek V4 Pro 0813
#172Worse than Luna but more expensive than Luna. Sticking with Luna without sending my data to Deepseek (China)
Unless you're Chinese, why would you care if they see your data? As an American, I'd much rather have my data kept outside the country than here where companies and the government have a lot more leverage over me.
Re: DeepSeek V4 Pro 0813
#173- https://api-docs.deepseek.com/
- https://x.com/ChrisGPT/status/2087572834650407024/photo/1 (officially posted on WeChat, this is just one of many reposts)
Re: DeepSeek V4 Pro 0813
#174Still behind Kimi-K3 in almost half of the benchmarks
Re: DeepSeek V4 Pro 0813
#175Tested both DS v4 pro 0813 and Grok 4.6 (all from openrouter) on Codex cli. Worked on a same new feature development on my project. Deepseek 4 pro: Worked for 12m 02s - cost $0.12 - has bug. Grok 4.6: Worked for 3m 18s - cost $ 1.41 - no bug.
Re: DeepSeek V4 Pro 0813
#176Tested both DS v4 pro 0813 and Grok 4.6 (all from openrouter) on Codex cli. Worked on a same new feature development on my project. Deepseek 4 pro: Worked for 12m 02s - cost $0.12 - has bug. Grok 4.6: Worked for 3m 18s - cost $ 1.41 - no bug.
I thought it was impossible to downvote posts?
User Posts can be downvoted but you need over 500 karma to have access to the downvote button. A Submission can not be downvoted.
Re: DeepSeek V4 Pro 0813
#177Earlier quoted context omitted.
The Deepseek official API is good with excellent caching. But their privacy policy is unusually bad - they can train off your prompts and completions.
Use another provider from OpenRouter. I really don’t care if they train off my prompts.
Re: DeepSeek V4 Pro 0813
#178Earlier quoted context omitted.
I don't think so. I specifically kept this PR to test model capabilities, and I've already tested a bunch of models. Current test results show that the more advanced the model is, the easier it passes. For example, GPT-5.5 Medium fails the test(has bug), but High passed.
They are causal autoregressive models, the output is sensitive even to the implementation nuances in inference. Even 1 token that's badly selected could throw off the entire answer.
Re: DeepSeek V4 Pro 0813
#179Just tested through openrouter.. gave exactly same task.. the task was to scan existing repo, and generate a single docker-compose file to deploy behind a caddy server, where certain port ranges are already used, the service demands widlcard certificates to be provisioned from outside, and postgre needs to be built-in one... Tested this model, and gpt-5.6-terra-high. Results: this one had few issues. terra: none. The…
Terra has not been able to do any of the technical tasks I've asked of it correctly. I'm surprised others get use out of it. Anything below Sol high tends to give me mostly unreliable results. I'm using codex as my main harness but maybe it performs better with a different one.
I think the key is to give them a nice assortment of self-verification tools, an AGENTS.md or reference document that they're encouraged to routinely check, and asking the planner to be thorough with the ACs but give the model some space.
The planner routinely finds issues with the worker's output, but that's what it is for.
Re: DeepSeek V4 Pro 0813
#180https://api-docs.deepseek.com/quick_start/pricing/ Competitive with opus 4.8 but weaker than sol or fable. About 20x cheaper.