Live data from Hacker News

DeepSeek V4 Pro 0813

openrouter.ai

381–390 of 493 posts

Re: DeepSeek V4 Pro 0813

#381

What I care about is whether the model is capable of the tasks I give it at the lowest cost. Right now I'm using Kimi-K3/GLM-5.2/Minimax. Sonnet is great but I burn through the tokens too fast. Opus 5 set to max is amazing and more intelligent than all of us. .998 of the time I don't need that kind of intelligence. I just need the job done.

Opus 5 fucking sucks to talk to and read compared to 5.6 Sol though. I’m fully done with Claude models until they figure this out

Re: DeepSeek V4 Pro 0813

#382
post #371
post #136

Nice bicycle chain, the little basket with a fish didn't show up in the right place: https://tools.simonwillison.net/markdown-svg-renderer#url=ht...

For a while now, I've found pelican rendering to be an unreliable metric for LLM ability - and most people know it. Yet, somehow it gets upvoted to the very top of every new model discussion.

becuase most people don't care whether it's accurate, as long as it looks right and is funny...

Re: DeepSeek V4 Pro 0813

#383

What I care about is whether the model is capable of the tasks I give it at the lowest cost. Right now I'm using Kimi-K3/GLM-5.2/Minimax. Sonnet is great but I burn through the tokens too fast. Opus 5 set to max is amazing and more intelligent than all of us. .998 of the time I don't need that kind of intelligence. I just need the job done.

Opus 5 fucking sucks to talk to and read compared to 5.6 Sol though. I’m fully done with Claude models until they figure this out

My god the Jargon is so hard to parse. Everytime i resort to cursing it , it understand. Even putting `ASD-STE100` or simplified english in the claude.md doesn't work. It gets the job done but is an anti social asshole

Re: DeepSeek V4 Pro 0813

#384
post #173

Why does this link to OpenRouter, which has no useful information on its own? Linking to the official API or the benchmarks would make more sense: - https://api-docs.deepseek.com/ - https://x.com/ChrisGPT/status/2087572834650407024/photo/1 (officially posted on WeChat, this is just one of many reposts)

I don’t know about you but I find the information about prices, effective price (weighted average), providers and performance, benchmarks (down bottom) very useful. With openrouter I can even test it right away and compare with other models (use the chat functions).

As of now, there is only a single provider for this model and that's the official DeepSeek API.

When GPT 6 comes out, would you expect the top thread to link to OpenRouter?

Re: DeepSeek V4 Pro 0813

#385

Even though cost-per-token is low, Deepseek v4 tends to burn an immense number of tokens to accomplish tasks.

I can go for days on end without topping up my deepseek account. When it can’t solve a problem I switch to GPT and have to top up in real time.

Re: DeepSeek V4 Pro 0813

#386
post #58
post #33

Earlier quoted context omitted.

i dont see any price increase there... what am i missing?

It's a big confusion, some[0] say an email was sent about significant price increase, personal I haven't seen anything official [0] https://finance.yahoo.com/technology/ai/articles/deepseek-pl...

Also got the email. It warned of a future large price increase, and to carefully watch usage.

I read it as a "hey we will make stuff more expensive, don't miss it"

Re: DeepSeek V4 Pro 0813

#387
Deepseek V4 Pro 0813 is the most unreliable model I have tried, it works on pass@3 shockingly well you can get it to match Sol or Fable perhaps in task done, but it's horrendous at pass@1 very prone to going wrong and doing horribly at most benches.

I am not sure what it is buy I suspect it might be GRPO.

Re: DeepSeek V4 Pro 0813

#388
Based on my experience so far, compared to previous models, DeepSeek V4 Pro achieves results equal to or even better than before, but at a lower cost.

Re: DeepSeek V4 Pro 0813

#390
post #375

So flash is 52 points on artificial analysis, and pro is 53

This was a disappointed to me. Why would I use pro over flash now? Is there some area where the difference is significant?

The idea is, I believe, that the Pro model is a larger model (more parameters, or less quantization) in general. What implication that has, I couldn't tell you.

For tasks like pondering on something, reviewing code, etc. I use Pro, just because it feels like the right model for that.

Post reply on HN