I really enjoyed this article. I think the author is precisely right and I've been saying this for a long time. There's a ton of extremely interesting low hanging fruit that can vastly improve the effectiveness of even currently existing models hiding in how we design our agent harnesses; enough to — at least until we hit diminishing returns — make as much or more of a difference than training new models! I think one…
Improving 15 LLMs at Coding in One Afternoon. Only the Harness Changed
71–80 of 318 posts
Re: Improving 15 LLMs at Coding in One Afternoon. Only the Harness Changed
#72I really enjoyed this article. I think the author is precisely right and I've been saying this for a long time. There's a ton of extremely interesting low hanging fruit that can vastly improve the effectiveness of even currently existing models hiding in how we design our agent harnesses; enough to — at least until we hit diminishing returns — make as much or more of a difference than training new models! I think one…
[flagged]
Re: Improving 15 LLMs at Coding in One Afternoon. Only the Harness Changed
#73I really enjoyed this article. I think the author is precisely right and I've been saying this for a long time. There's a ton of extremely interesting low hanging fruit that can vastly improve the effectiveness of even currently existing models hiding in how we design our agent harnesses; enough to — at least until we hit diminishing returns — make as much or more of a difference than training new models! I think one…
Also, yes, I'm aware that I use a lot of "its not just X, its Y." I promise you this comment is entirely human written. I'm just really tired and tend to rely on more wrote rhetorical tropes when I am. Believe me, I wrote like this long before LLMs were a thing.
Re: Improving 15 LLMs at Coding in One Afternoon. Only the Harness Changed
#74My personal notes (not the author): have been way faster performance wise which is honestly the biggest improvement over correctless. I've posted https://github.com/can1357/oh-my-pi before, but didn't seem to gain traction. It's a great little agent.
I've just started messing around with pi, but haven't fully dug in yet. How would you compare oh-my-pi? I see it has a lot of other bells and whistles built in. Are they portable bit by bit back to pi, or is there enough differences that they can't? how about normal pi extensions, can they be used in omp? Some of the stuff definitely looks interesting.
Re: Improving 15 LLMs at Coding in One Afternoon. Only the Harness Changed
#75Re: Improving 15 LLMs at Coding in One Afternoon. Only the Harness Changed
#76I see a lot of evidence to the contrary though. Anyone know what the underlying issue here is?
Re: Improving 15 LLMs at Coding in One Afternoon. Only the Harness Changed
#77I really enjoyed this article. I think the author is precisely right and I've been saying this for a long time. There's a ton of extremely interesting low hanging fruit that can vastly improve the effectiveness of even currently existing models hiding in how we design our agent harnesses; enough to — at least until we hit diminishing returns — make as much or more of a difference than training new models! I think one…
[flagged]
Re: Improving 15 LLMs at Coding in One Afternoon. Only the Harness Changed
#78Is it possible that burning extra tokens is the point, since they get paid more?
Re: Improving 15 LLMs at Coding in One Afternoon. Only the Harness Changed
#79Re: Improving 15 LLMs at Coding in One Afternoon. Only the Harness Changed
#80I really enjoyed this article. I think the author is precisely right and I've been saying this for a long time. There's a ton of extremely interesting low hanging fruit that can vastly improve the effectiveness of even currently existing models hiding in how we design our agent harnesses; enough to — at least until we hit diminishing returns — make as much or more of a difference than training new models! I think one…
[flagged]