Show HN: FrontierHarness Eval – 9 harness, same model, cost per pass varies 17x
31–40 of 67 posts
Re: Show HN: FrontierHarness Eval – 9 harness, same model, cost per pass varies 17x
#32People always talk about how Cursor harness has some secret sauce; would like to see how that one stacks up.
Thanks for making this and filling a real gap!
Re: Show HN: FrontierHarness Eval – 9 harness, same model, cost per pass varies 17x
#33One issue I see with including something like Pi is that it's intentionally bare bones. I don't think anyone uses Pi without some basic custom extensions (subagents, check lists, etc), so this benchmark may not be representative of a realistic setup. Would be interesting to see some sort of ablation test too. E.g. what parts of Codex contribute most to the perf, and can they be recreated more minimally in something l…
It's even weirder given that another of the test cases is essentially "We put parts on the frame (OhMyPi) and it went faster!"
Re: Show HN: FrontierHarness Eval – 9 harness, same model, cost per pass varies 17x
#34One issue I see with including something like Pi is that it's intentionally bare bones. I don't think anyone uses Pi without some basic custom extensions (subagents, check lists, etc), so this benchmark may not be representative of a realistic setup. Would be interesting to see some sort of ablation test too. E.g. what parts of Codex contribute most to the perf, and can they be recreated more minimally in something l…
Re: Show HN: FrontierHarness Eval – 9 harness, same model, cost per pass varies 17x
#35Re: Show HN: FrontierHarness Eval – 9 harness, same model, cost per pass varies 17x
#36Forgive my stupidity but how do you run Claude Code with non-Anthropic models?
Re: Show HN: FrontierHarness Eval – 9 harness, same model, cost per pass varies 17x
#37Do you have instructions on how to run a custom harness against this? There are none on the linked page. I want to run Dirac ( https://github.com/dirac-run/dirac ).
Blog Post: https://runta.com/blog/introducing-frontierharness-eval/
The tasks and methodology: https://github.com/runta-dev/frontier-harness-eval
Re: Show HN: FrontierHarness Eval – 9 harness, same model, cost per pass varies 17x
#38Re: Show HN: FrontierHarness Eval – 9 harness, same model, cost per pass varies 17x
#39Is Github Copilot (integrated in VS Code) not a thing? I use it all the time and don't know what more I could wish for. Disclaimer: I don't let agents run on huge tasks for hours. Almost all tasks I give them are done in under 30 min.
Re: Show HN: FrontierHarness Eval – 9 harness, same model, cost per pass varies 17x
#40One issue I see with including something like Pi is that it's intentionally bare bones. I don't think anyone uses Pi without some basic custom extensions (subagents, check lists, etc), so this benchmark may not be representative of a realistic setup. Would be interesting to see some sort of ablation test too. E.g. what parts of Codex contribute most to the perf, and can they be recreated more minimally in something l…
It does seem a bit odd to say "We tested all of these bicycles with the same rider" when one of the test cases is actually a bare high-end frame with no components on it. It's even weirder given that another of the test cases is essentially "We put parts on the frame (OhMyPi) and it went faster!"
With the bicycle, some may prefer comfort, others speed, others offroad, etc., so it would not be obvious which one is "best".