DeepSeek V4 Pro 0813
391–400 of 493 posts
Re: DeepSeek V4 Pro 0813
#392Just tested through openrouter.. gave exactly same task.. the task was to scan existing repo, and generate a single docker-compose file to deploy behind a caddy server, where certain port ranges are already used, the service demands widlcard certificates to be provisioned from outside, and postgre needs to be built-in one... Tested this model, and gpt-5.6-terra-high. Results: this one had few issues. terra: none. The…
Terra has not been able to do any of the technical tasks I've asked of it correctly. I'm surprised others get use out of it. Anything below Sol high tends to give me mostly unreliable results. I'm using codex as my main harness but maybe it performs better with a different one.
Re: DeepSeek V4 Pro 0813
#393Nice bicycle chain, the little basket with a fish didn't show up in the right place: https://tools.simonwillison.net/markdown-svg-renderer#url=ht...
Re: DeepSeek V4 Pro 0813
#394Earlier quoted context omitted.
Can we have a conversation about subscription plans for a minute? I don't mean to hype up the US AI firms, but if a ChatGPT $200/m subscription can get you $16,000 in effective API costs, doesn't effectively every model get destroyed by the subsidized Claude/ChatGPT models? Both in price and intelligence.
That is the problem currently the subscription plans are being subsidized by VC money and token buyers. When Open weight get good enough token buyers build their own servers instead of buying tokens then no one to subsidize the subscriptions
Cycle forward to Fable 7, Kimi 5, GPT 7 a couple years out. Forget about it unless you own a datacenter.
Re: DeepSeek V4 Pro 0813
#395Earlier quoted context omitted.
This has not been my experience. Generally I do pin to 1 provider, or 1 provider with a couple fallbacks (especially with deepseek - most providers are 10x the cached token price compared to deepseek themselves), but even when I don't I still usually see 99%+ cache hit percentage. Specifically using pi with various ad-hoc customisations (that I was careful not to break prompt caching with).
then what is the point of using operouter for this model? Just use the deepseek API and save the 5% fee on top of the better caching rate.
Re: DeepSeek V4 Pro 0813
#396Earlier quoted context omitted.
Right now, sol-xhigh is my favorite model. I feel that Opus 5 is dumber than 4.8. Fable is too expensive to do anything (limit of $50, started a prompt at $25, ended up at $75, is bullshit, but at least it's "free credits"). DeepSeek is okay for random API-based stuff, as it's cheap. Local open models running on a 5090 are hit or miss. I feel that most GGUFs/quants are awful...
Opus 5 degrades to word salad. I wonder if it is because of watermarking.
Re: DeepSeek V4 Pro 0813
#397Nice bicycle chain, the little basket with a fish didn't show up in the right place: https://tools.simonwillison.net/markdown-svg-renderer#url=ht...
For a while now, I've found pelican rendering to be an unreliable metric for LLM ability - and most people know it. Yet, somehow it gets upvoted to the very top of every new model discussion.
Re: DeepSeek V4 Pro 0813
#398Earlier quoted context omitted.
That is the problem currently the subscription plans are being subsidized by VC money and token buyers. When Open weight get good enough token buyers build their own servers instead of buying tokens then no one to subsidize the subscriptions
99.9%+ of the tech worker population will never be able to build their own servers to run future frontier models. Kimi 3 is an indication of what's coming. These models will keep getting drastically larger. The hardware isn't getting cheaper anytime soon (no matter what China does; that goes for memory and GPUs). Cycle forward to Fable 7, Kimi 5, GPT 7 a couple years out. Forget about it unless you own a datacenter.
A single local user can run frontier models slowly on a 24/7 basis, which drops hardware requirements by orders of magnitude compared to a datacenter setup for just-in-time inference. This is not a real alternative to subsidized subscriptions at present, but it's a great insurance policy against future VC-driven rug pulls.
Re: DeepSeek V4 Pro 0813
#399Earlier quoted context omitted.
Opus 5 degrades to word salad. I wonder if it is because of watermarking.
Opus 5 doesn't really even speak coherent English. I'm not sure what's going on, but it can't explain anything. It still does an excellent job with code and writing tests and code review and creating and completing a plan, and it seems to be able to understand English instructions, but it sure as hell can't explain what it did or how to use the code it wrote. That was true before they announced the watermarking, I'd…
Re: DeepSeek V4 Pro 0813
#400Nice bicycle chain, the little basket with a fish didn't show up in the right place: https://tools.simonwillison.net/markdown-svg-renderer#url=ht...