Live data from Hacker News

DeepSeek V4 Pro 0813

openrouter.ai

441–450 of 493 posts

Re: DeepSeek V4 Pro 0813

#441

I had high expectations for V4 Pro, especially since DeepSeek V4 Flash 0731 performed so well compared with other Flash models. What a letdown.

Wait, what are you disappointed by? Seems like a significant jump in performance, and it competes handily with other models.

Re: DeepSeek V4 Pro 0813

#442

Deepseek V4 Pro 0813 is the most unreliable model I have tried, it works on pass@3 shockingly well you can get it to match Sol or Fable perhaps in task done, but it's horrendous at pass@1 very prone to going wrong and doing horribly at most benches. I am not sure what it is buy I suspect it might be GRPO.

Could you try setting the temperature very low e.g. 0.0?

Re: DeepSeek V4 Pro 0813

#445
post #329

Earlier quoted context omitted.

It is a pain from openRouter if you don't define your providers correctly, but for DeepSeek, surely not- the weights aren't released yet and there's only one provider, DeepSeek.

With Deepseek as the provider, there's no issue of course, but that means you don't filter providers for data retention, and you could also choose direct API use with them at that point.

Fallback are still very useful and won't poison much your cache hits too much if the provider is down anyway.

Re: DeepSeek V4 Pro 0813

#446
Deepseek V4 Flash 0731 was such a massive jump in capability for such a small model (and price), that I'm a bit disappointed by this release.

I keep my agents on tight leashes, using them very interactively for bouncing off ideas, architecture, and then writing code (especially prototyping) and Flash has been crushing everything I ever needed it to do.

Maybe my ambitions are too tame compared to people needing Fable / Sol grade models, but I'm probably staying on Flash and not moving on to Pro for the foreseeable future.

Re: DeepSeek V4 Pro 0813

#447

Earlier quoted context omitted.

Right now, sol-xhigh is my favorite model. I feel that Opus 5 is dumber than 4.8. Fable is too expensive to do anything (limit of $50, started a prompt at $25, ended up at $75, is bullshit, but at least it's "free credits"). DeepSeek is okay for random API-based stuff, as it's cheap. Local open models running on a 5090 are hit or miss. I feel that most GGUFs/quants are awful...

Opus 5 degrades to word salad. I wonder if it is because of watermarking.

Ah, so it's not only me :-D

I think it could be the watermarking, but at this point they might be deliberately complicating the prose so that we ask clarifying questions and that leads to more token spend.

Re: DeepSeek V4 Pro 0813

#449
OpenRouter is a place with zero support.

Suddenly get a big debt on your account with nobody to respond.

As an early adopter of OpenRouter, I'm afraid they are in shambles.

Re: DeepSeek V4 Pro 0813

#450
post #408

Been testing this on hobby project https://github.com/arj03/seedkernel/ . Latest flash was a big step up. Pro feels really slow compared. Claude opus is still better day this level.

That said. It is really good at security review at max settings. Just ran one, it came to 0.15$
Post reply on HN