Live data from Hacker News

DeepSeek V4 Pro 0813

openrouter.ai

291–300 of 493 posts

Re: DeepSeek V4 Pro 0813

#291
post #223

Earlier quoted context omitted.

[flagged]

If I live my life on the basis that some people don't share my sense of humor, and hence I should avoid doing anything funny that might be misunderstood, my life will be a lot less fun.

You're doing great. Don't let insanely low-effort (negative-effort, as in making others dumber rather than having no effect?) comments like from the above throwaway affect your actions.

Re: DeepSeek V4 Pro 0813

#292

What I care about is whether the model is capable of the tasks I give it at the lowest cost. Right now I'm using Kimi-K3/GLM-5.2/Minimax. Sonnet is great but I burn through the tokens too fast. Opus 5 set to max is amazing and more intelligent than all of us. .998 of the time I don't need that kind of intelligence. I just need the job done.

Check out https://unbiased.ai

(Disclaimer: I’m a co-founder)

Re: DeepSeek V4 Pro 0813

#293
post #289
post #173

Why does this link to OpenRouter, which has no useful information on its own? Linking to the official API or the benchmarks would make more sense: - https://api-docs.deepseek.com/ - https://x.com/ChrisGPT/status/2087572834650407024/photo/1 (officially posted on WeChat, this is just one of many reposts)

There's no new page for this model. Hackernews didn't allow the same link be posted twice.

Bogus query params could work maybe

Re: DeepSeek V4 Pro 0813

#294

Earlier quoted context omitted.

When it comes to quality of outcome, since at least Feburary, the harness has almost equal, if not more weight than the model itself. It's no longer "which model is the best?" it's "which model + harness is the best?" I get drastically different tool call failure rates using Claude SDK vs OpenCode using Qwen 3.6 models

Every other week it's a new "X didn't matter, until Y date" without any hard quantitative claims. It's crazy how over the past years a field originating from math ends up succumbing to subjective feels.

It's a forum, not an engineering conference, but here is a demonstrated 10 point difference between two top harnesses anyways: https://artificialanalysis.ai/agents/coding-agents#harness-c...

Re: DeepSeek V4 Pro 0813

#295
post #136

Nice bicycle chain, the little basket with a fish didn't show up in the right place: https://tools.simonwillison.net/markdown-svg-renderer#url=ht...

On either side of the front wheel is a perfectly reasonable place to carry cargo. I think I'd have taken more issue with the spokes, or at least that's what stood out to me. The chain is indeed nice, however.

Re: DeepSeek V4 Pro 0813

#296
post #173

Why does this link to OpenRouter, which has no useful information on its own? Linking to the official API or the benchmarks would make more sense: - https://api-docs.deepseek.com/ - https://x.com/ChrisGPT/status/2087572834650407024/photo/1 (officially posted on WeChat, this is just one of many reposts)

Moreover, OpenRouter is NOT Open Source, fair source, source available, etc. It's a proprietary cloud service that got first place in the API aggregation distribution game. Link to DeepSeek!

Open doesn't always refer to the code. Just like their previous project, it refers to an open marketplace where anybody can sign up to sell access to models.

But it'd still be nice to post to wait an extra minute to find some other page/new url from deepseek for it instead of posting that it exists somewhere.

Re: DeepSeek V4 Pro 0813

#297

Earlier quoted context omitted.

IME I can't trust it to write it's own plans from a spec, but if I give it a detailed execution plan written by Opus, it's fast and cheap (if chatty) in executing it.

Interesting. I use Flash for making the plans and GPT for execution.

Depending on the language you're writing in and the problem domain, the smaller models can do dramatically better or worse.

I suspect in the future we'll see language-specific small models. "Coding" is still pretty broad as an activity. It'd be nice to be able to load up a model specific to, say, class-based Python and run it on-device.

Re: DeepSeek V4 Pro 0813

#298
post #223

Earlier quoted context omitted.

[flagged]

If I live my life on the basis that some people don't share my sense of humor, and hence I should avoid doing anything funny that might be misunderstood, my life will be a lot less fun.

I’m always excited to see “simonw” on my screen.

More likely than not to be interesting & accessible for those of us outside e.g. compsci.

Hope you stay happy so you stay nerdy and keep sharing both with us!

Re: DeepSeek V4 Pro 0813

#299

Earlier quoted context omitted.

How do you define intelligence? I encounter that kind of sentiment all too often, and I have to assume we go by wildly different understanding of what that might entail.

If you read Opus 5's output, it is beyond the comprehension of virtually all engineers and developers. That is what I mean by intelligence. Math, science, and engineering are all contained in one model. We may be experts in one field. The model is an expert in everything that humans know.

Careful, you may have a bit of psychosis. They are very, very far from incomprehensible, and also very far from the top at least of my field. The best in my field are produce far higher quality results, and I think that's true for all fields. It's just an incredibly good 85% quality machine that experts all use because they can guide it to be up to their quality faster than doing it themselves.

Re: DeepSeek V4 Pro 0813

#300

Earlier quoted context omitted.

Their leaks would confirm this sort of attitude. They're not trying to become the top player or anything like that - just working to play their part in pushing LLM tech forward and going from there. It was quite refreshing from the 'here's how we're going to dominate the world' nonsense. It's undoubtedly the same attitude that just lets them shrug and cancel the fund raising round after the leaks came from said fundi…

The founder of DS's stated goal is to get to AGI. He thinks this is the path to get there. Kind of interesting, when compared to the hubris from American frontier labs.

When one needs money, an infinitely remote goal is the best cause.
Post reply on HN