Improving 15 LLMs at Coding in One Afternoon. Only the Harness Changed
91–100 of 318 posts
Re: Improving 15 LLMs at Coding in One Afternoon. Only the Harness Changed
#92I really enjoyed this article. I think the author is precisely right and I've been saying this for a long time. There's a ton of extremely interesting low hanging fruit that can vastly improve the effectiveness of even currently existing models hiding in how we design our agent harnesses; enough to — at least until we hit diminishing returns — make as much or more of a difference than training new models! I think one…
OpenAI used early versions of GPT-5.3-Codex to: debug its own training process, manage its deployment and scaling and diagnose test results and evaluation data.
Claude Code have shipped 22 PRs in a single day and 27 the day before, with 100% of the code in each PR generated entirely by Claude Code.
Re: Improving 15 LLMs at Coding in One Afternoon. Only the Harness Changed
#93I really enjoyed this article. I think the author is precisely right and I've been saying this for a long time. There's a ton of extremely interesting low hanging fruit that can vastly improve the effectiveness of even currently existing models hiding in how we design our agent harnesses; enough to — at least until we hit diminishing returns — make as much or more of a difference than training new models! I think one…
Re: Improving 15 LLMs at Coding in One Afternoon. Only the Harness Changed
#94The harness is the model "body", it's weight the cognition. Like in nature they develop together and the iteration of natural selection works at both. If smaller labs (Zai, Moonshot, deepseek, mistral..) get together and embrace a harness, like opencode for example, as a consortium just by the power of "evolution across different environments" they might hit jackpot earlier than bigger labs.
Someone has to do the baseline training, development, and innovation. it can't be clones all the way down
Re: Improving 15 LLMs at Coding in One Afternoon. Only the Harness Changed
#95I really enjoyed this article. I think the author is precisely right and I've been saying this for a long time. There's a ton of extremely interesting low hanging fruit that can vastly improve the effectiveness of even currently existing models hiding in how we design our agent harnesses; enough to — at least until we hit diminishing returns — make as much or more of a difference than training new models! I think one…
[flagged]
Re: Improving 15 LLMs at Coding in One Afternoon. Only the Harness Changed
#96"You're absolutely right!"
At this point I'd take a contract with Anthropic to have Claude code pick better tooling.
Re: Improving 15 LLMs at Coding in One Afternoon. Only the Harness Changed
#97I really enjoyed this article. I think the author is precisely right and I've been saying this for a long time. There's a ton of extremely interesting low hanging fruit that can vastly improve the effectiveness of even currently existing models hiding in how we design our agent harnesses; enough to — at least until we hit diminishing returns — make as much or more of a difference than training new models! I think one…
[flagged]
Re: Improving 15 LLMs at Coding in One Afternoon. Only the Harness Changed
#98Re: Improving 15 LLMs at Coding in One Afternoon. Only the Harness Changed
#99Re: Improving 15 LLMs at Coding in One Afternoon. Only the Harness Changed
#100Problem is, replace has been around for so long, most LLMs are tuned for it now