Interesting. Did you give the agents any skills for playing civ? If not, are you planning to?
Show HN: CivBench a long-horizon AI benchmark for multi-agent games
11–20 of 26 posts
Re: Show HN: CivBench a long-horizon AI benchmark for multi-agent games
#12Have you tried playing the agents yourself? Do they crush human competition?
Re: Show HN: CivBench a long-horizon AI benchmark for multi-agent games
#13This is great. I think leaderboards based on static evals will be mostly irrelevant within a year. Continuous benchmarks like this are the only way to get signal on frontier models You mention Opus 4.6 cost $1200 in one match, how do you plan to benchmark economic efficiency? Looking at a performance vs. cost trade-off you might say a model that plays 80% as well at 1% of the cost is more impressive than the 'top' mo…
In the leaderboards part of the page I'll be autopopulating the token cost of the model as a metric to evaluate on
Re: Show HN: CivBench a long-horizon AI benchmark for multi-agent games
#14Re: Show HN: CivBench a long-horizon AI benchmark for multi-agent games
#15Re: Show HN: CivBench a long-horizon AI benchmark for multi-agent games
#16Re: Show HN: CivBench a long-horizon AI benchmark for multi-agent games
#17Re: Show HN: CivBench a long-horizon AI benchmark for multi-agent games
#18Re: Show HN: CivBench a long-horizon AI benchmark for multi-agent games
#19Incredible and important product. Necessary for developers, users, and industries that want to use agents. Can’t wait to see how it’ll grow
Re: Show HN: CivBench a long-horizon AI benchmark for multi-agent games
#20This is an amazing eval metric that no one thought about! such a creative idea. Have you thought of other games? how different it is from chess?