Live data from Hacker News

Show HN: FrontierHarness Eval – 9 harness, same model, cost per pass varies 17x

frontierharness.org

21–30 of 67 posts

Re: Show HN: FrontierHarness Eval – 9 harness, same model, cost per pass varies 17x

#21

With all the hype around the latest Gemini 3.8 Flash/Cyber release, will Antigravity CLI [1] be supported? 1/ https://antigravity.google/product/antigravity-cli

agy has fewer users than grok, both are ~1% based on some surveys I've seen, not every harness needs to be evaluated

Re: Show HN: FrontierHarness Eval – 9 harness, same model, cost per pass varies 17x

#22
post #12

Forgive my stupidity but how do you run Claude Code with non-Anthropic models?

Claude Code supports the base URL env var so you could tell it to talk with any LLM API endpoint that receives the Anthropic style request format, e.g. DeepSeek.

The full list for the curious

https://code.claude.com/docs/en/env-vars#variables

Re: Show HN: FrontierHarness Eval – 9 harness, same model, cost per pass varies 17x

#25
post #17

Good start, but needs to be harness x model to be useful. 3 top harnesses x 3 top models would be more interesting.

Version 1.1 will add more harnesses and evaluate a broader set of models. Rather than holding the model fixed, we plan to evaluate the full harness × model matrix to reveal interaction effects (which harnesses work best with which models) and produce a harness-model compatibility matrix.

Re: Show HN: FrontierHarness Eval – 9 harness, same model, cost per pass varies 17x

#29
One issue I see with including something like Pi is that it's intentionally bare bones. I don't think anyone uses Pi without some basic custom extensions (subagents, check lists, etc), so this benchmark may not be representative of a realistic setup.

Would be interesting to see some sort of ablation test too. E.g. what parts of Codex contribute most to the perf, and can they be recreated more minimally in something like Pi.

Post reply on HN