Live data from Hacker News

DeepSeek V4 Pro 0813

openrouter.ai

401–410 of 493 posts

Re: DeepSeek V4 Pro 0813

#401
post #371
post #136

Nice bicycle chain, the little basket with a fish didn't show up in the right place: https://tools.simonwillison.net/markdown-svg-renderer#url=ht...

For a while now, I've found pelican rendering to be an unreliable metric for LLM ability - and most people know it. Yet, somehow it gets upvoted to the very top of every new model discussion.

I prefer the Browser OS test

Re: DeepSeek V4 Pro 0813

#402
post #67

Earlier quoted context omitted.

Their leaks would confirm this sort of attitude. They're not trying to become the top player or anything like that - just working to play their part in pushing LLM tech forward and going from there. It was quite refreshing from the 'here's how we're going to dominate the world' nonsense. It's undoubtedly the same attitude that just lets them shrug and cancel the fund raising round after the leaks came from said fundi…

Benefits of having a well performing hedge fund funding DeepSeek. IIRC, Demis attempted to start a fund inside DeepMind but it was killed off. In an alternative world where he manages to pull that off, perhaps DeepMind would still be independent with Demis at the helm.

rumor is that is what Ilya has done at SSI.

Re: DeepSeek V4 Pro 0813

#403
post #371

Earlier quoted context omitted.

For a while now, I've found pelican rendering to be an unreliable metric for LLM ability - and most people know it. Yet, somehow it gets upvoted to the very top of every new model discussion.

becuase most people don't care whether it's accurate, as long as it looks right and is funny...

It doesn't look right at all.

Re: DeepSeek V4 Pro 0813

#404
I’ve found Flash 0731 to be pretty great recently. I do feel like a lot of the models are pretty close in terms of ability. I often run code through multiple different models _and_ harnesses for code reviews and they all typically find the same things as each other.

Re: DeepSeek V4 Pro 0813

#405

Earlier quoted context omitted.

Terra has not been able to do any of the technical tasks I've asked of it correctly. I'm surprised others get use out of it. Anything below Sol high tends to give me mostly unreliable results. I'm using codex as my main harness but maybe it performs better with a different one.

Terra is great. It's wild how different our experiences are. Install the Superpowers plugin. Behold.

In my experience, Superpowers has begun to massively slow down the capable models at this point. The skills they add are incredibly bloated and just get you worse results nowadays, tbh.

Re: DeepSeek V4 Pro 0813

#409

Earlier quoted context omitted.

> the harness has almost equal, if not more weight than the model itself This feels like a horrible failing of the models to generalize, then - both basic and intermediate tasks should be possible to do with Claude Code, OpenCode, Pi, ZCode, Kimi Code, Dirac and tbh any other mainstream or even slightly niche harness. Not doubting the claim itself, there's a reason why good benchmarks include the harness.

i think thats BS that harness has equal weight. most of intellegice is still coming from training data not from RL. so how is 'coevolved harness' equal weight.

> most of intellegice is still coming from training data not from RL.

For coding specifically, I'm not sure this is still true. Given the heavy use of RL to improve coding performance, I'd expect the harness to be important as it defines what tools the model is rewarded for using.

Re: DeepSeek V4 Pro 0813

#410
post #136

Nice bicycle chain, the little basket with a fish didn't show up in the right place: https://tools.simonwillison.net/markdown-svg-renderer#url=ht...

You don’t need the best model in 99% of cases…

You don’t need Fable or Sol to execute tasks. However, you need them to supervise and plan. Like, Luna is cheap and is at DSV4F level, but it’s not capable of advanced reasoning.
Post reply on HN