Nice bicycle chain, the little basket with a fish didn't show up in the right place: https://tools.simonwillison.net/markdown-svg-renderer#url=ht...
For a while now, I've found pelican rendering to be an unreliable metric for LLM ability - and most people know it. Yet, somehow it gets upvoted to the very top of every new model discussion.
DeepSeek V4 Pro 0813
401–410 of 493 posts
Re: DeepSeek V4 Pro 0813
#402Earlier quoted context omitted.
Their leaks would confirm this sort of attitude. They're not trying to become the top player or anything like that - just working to play their part in pushing LLM tech forward and going from there. It was quite refreshing from the 'here's how we're going to dominate the world' nonsense. It's undoubtedly the same attitude that just lets them shrug and cancel the fund raising round after the leaks came from said fundi…
Benefits of having a well performing hedge fund funding DeepSeek. IIRC, Demis attempted to start a fund inside DeepMind but it was killed off. In an alternative world where he manages to pull that off, perhaps DeepMind would still be independent with Demis at the helm.
Re: DeepSeek V4 Pro 0813
#403Earlier quoted context omitted.
For a while now, I've found pelican rendering to be an unreliable metric for LLM ability - and most people know it. Yet, somehow it gets upvoted to the very top of every new model discussion.
becuase most people don't care whether it's accurate, as long as it looks right and is funny...
Re: DeepSeek V4 Pro 0813
#404Re: DeepSeek V4 Pro 0813
#405Earlier quoted context omitted.
Terra has not been able to do any of the technical tasks I've asked of it correctly. I'm surprised others get use out of it. Anything below Sol high tends to give me mostly unreliable results. I'm using codex as my main harness but maybe it performs better with a different one.
Terra is great. It's wild how different our experiences are. Install the Superpowers plugin. Behold.
Re: DeepSeek V4 Pro 0813
#406So flash is 52 points on artificial analysis, and pro is 53
Re: DeepSeek V4 Pro 0813
#407Re: DeepSeek V4 Pro 0813
#408Re: DeepSeek V4 Pro 0813
#409Earlier quoted context omitted.
> the harness has almost equal, if not more weight than the model itself This feels like a horrible failing of the models to generalize, then - both basic and intermediate tasks should be possible to do with Claude Code, OpenCode, Pi, ZCode, Kimi Code, Dirac and tbh any other mainstream or even slightly niche harness. Not doubting the claim itself, there's a reason why good benchmarks include the harness.
i think thats BS that harness has equal weight. most of intellegice is still coming from training data not from RL. so how is 'coevolved harness' equal weight.
For coding specifically, I'm not sure this is still true. Given the heavy use of RL to improve coding performance, I'd expect the harness to be important as it defines what tools the model is rewarded for using.
Re: DeepSeek V4 Pro 0813
#410Nice bicycle chain, the little basket with a fish didn't show up in the right place: https://tools.simonwillison.net/markdown-svg-renderer#url=ht...
You don’t need the best model in 99% of cases…