[flagged]
Computer Anthology: A continuously evolving benchmark family for AI agents
11–20 of 20 posts
Re: Computer Anthology: A continuously evolving benchmark family for AI agents
#12[flagged]
Re: Computer Anthology: A continuously evolving benchmark family for AI agents
#13[dead]
Re: Computer Anthology: A continuously evolving benchmark family for AI agents
#14[dead]
Re: Computer Anthology: A continuously evolving benchmark family for AI agents
#15The comparison between harnesses is very nice. Interesting to see that using a different harness can bump the performance of the model as much as a new version (e.g., GPT 5.5+Codex ~= GPT 5.6+Terminus, at lower cost)
That’s a nice discussion. Some people say that with current model capabilities, the real differentiator is the harness. What are the best harnesses you guys are using?
Re: Computer Anthology: A continuously evolving benchmark family for AI agents
#16good work!
Re: Computer Anthology: A continuously evolving benchmark family for AI agents
#17Interesting the idea of treating the benchmark as an evolving system rather than a static dataset.
Re: Computer Anthology: A continuously evolving benchmark family for AI agents
#18[dead]
Re: Computer Anthology: A continuously evolving benchmark family for AI agents
#19Impressive!
Re: Computer Anthology: A continuously evolving benchmark family for AI agents
#20[flagged]