Live data from Hacker News

DeepSeek V4 Pro 0813

openrouter.ai

431–440 of 493 posts

Re: DeepSeek V4 Pro 0813

#432
post #136

Nice bicycle chain, the little basket with a fish didn't show up in the right place: https://tools.simonwillison.net/markdown-svg-renderer#url=ht...

Wonder if we'll ever see optimizations for pelican riding a bicycle svg, make it in to model training runs.

Re: DeepSeek V4 Pro 0813

#433

Earlier quoted context omitted.

Terra has not been able to do any of the technical tasks I've asked of it correctly. I'm surprised others get use out of it. Anything below Sol high tends to give me mostly unreliable results. I'm using codex as my main harness but maybe it performs better with a different one.

the breadth and width of the universe of oneshot challenges are all arbitrary. It's unsurprising different workflows oneshot better than others. All the more reason to favor local models under your control, as once you find that sweet spot model, no one can change it, upgrade it, align it, take it down or otherwise harm the time investment you made it making it your own. I can't really believe no one understands, aft…

But I don’t want to manage a local model…

Re: DeepSeek V4 Pro 0813

#434

Earlier quoted context omitted.

50% cache hit is really low - in a standard agentic loop you should expect like 99%+ cache hit percentage (which should also lower that $12.50 to like a couple of $ for the same amount of tokens). If you're using a customised harness you should make sure you don't have something that's e.g. changing your system prompt on some requests or rewriting history - it can be tempting to do stuff like strip old thinking token…

In my experience, that's the OpenRouter tax. Even a session that does everything right to remain sticky ends up getting moved between providers on a few requests, which bills you the full context as input every time the switch happens. I assume it's done as load balancing/latency mitigation, but it's put me off of OpenRouter for my use cases (limited use, limited need for changing models).

You can set it up to always use the official provider.

Re: DeepSeek V4 Pro 0813

#435

Earlier quoted context omitted.

Can we have a conversation about subscription plans for a minute? I don't mean to hype up the US AI firms, but if a ChatGPT $200/m subscription can get you $16,000 in effective API costs, doesn't effectively every model get destroyed by the subsidized Claude/ChatGPT models? Both in price and intelligence.

Yeah pretty much. I spent half a billion in tokens one night on a huge refactor with DSFlash, cost $11. If I spent that every night it would be 3x my GPT subscription.

That seems too expensive to be honest. Did you do it with official deepseek api or a 3rd party provider? Because official has 10x cheaper cache reads than the rest. I've done similar sized chats for like $1

Re: DeepSeek V4 Pro 0813

#436

Earlier quoted context omitted.

50% cache hit is really low - in a standard agentic loop you should expect like 99%+ cache hit percentage (which should also lower that $12.50 to like a couple of $ for the same amount of tokens). If you're using a customised harness you should make sure you don't have something that's e.g. changing your system prompt on some requests or rewriting history - it can be tempting to do stuff like strip old thinking token…

In my experience, that's the OpenRouter tax. Even a session that does everything right to remain sticky ends up getting moved between providers on a few requests, which bills you the full context as input every time the switch happens. I assume it's done as load balancing/latency mitigation, but it's put me off of OpenRouter for my use cases (limited use, limited need for changing models).

Seems like pro 0813 is exclusively served by Deepseek themselves at the moment so I wouldn't say that's the case?

Re: DeepSeek V4 Pro 0813

#438

Did you guys read the fine print they plan to increase prices significantly in the future

Take a look at openrouter there are a few alternative deepseek model providers that follow an even cheaper pricing strategy.

Funnily even if deepseek themselves increase price 2-3x they are still more affordable.

Re: DeepSeek V4 Pro 0813

#439
post #366

The past DeepSeek models and now these new checkpoints score very badly on the ArtificialAnalysis AA-Omniscience and hallucination rate benchmarks. I wonder where that's from? Maybe they're overindexing on coding even more than others? I can't say I've noticed it in my (coding) usage so far, has anyone seen it make up potential root causes or other speculative stuff more than other models?

Yeah. I stopped ising deepseek v4 flashbecause it is awful (even worse than my local qwen3.6 35B model) at multilingual prose.

Working on a language related app makes me realize that all these supposed language models don't have many good language benchmarks

Re: DeepSeek V4 Pro 0813

#440
post #371
post #136

Nice bicycle chain, the little basket with a fish didn't show up in the right place: https://tools.simonwillison.net/markdown-svg-renderer#url=ht...

For a while now, I've found pelican rendering to be an unreliable metric for LLM ability - and most people know it. Yet, somehow it gets upvoted to the very top of every new model discussion.

At this point I want to see some human-drawn pelicans on bicycles. I suspect the LLMs aren't doing all that bad.
Post reply on HN