Live data from Hacker News

Show HN: CivBench a long-horizon AI benchmark for multi-agent games

clashai.live

11–20 of 26 posts

Re: Show HN: CivBench a long-horizon AI benchmark for multi-agent games

#11
post #7

Interesting. Did you give the agents any skills for playing civ? If not, are you planning to?

I want to! I think skills can add big performance gains here especially with smaller models. There's a lot of domain knowledge in games so distilling it into a "skill" may allow much smaller models to outcompete the large ones

Re: Show HN: CivBench a long-horizon AI benchmark for multi-agent games

#13
post #9

This is great. I think leaderboards based on static evals will be mostly irrelevant within a year. Continuous benchmarks like this are the only way to get signal on frontier models You mention Opus 4.6 cost $1200 in one match, how do you plan to benchmark economic efficiency? Looking at a performance vs. cost trade-off you might say a model that plays 80% as well at 1% of the cost is more impressive than the 'top' mo…

For a game that runs 4+ hours unfortunately it was configured to use too much reasoning/turn and larger context. Reducing the size helped lower the cost (still expensive).

In the leaderboards part of the page I'll be autopopulating the token cost of the model as a metric to evaluate on

Re: Show HN: CivBench a long-horizon AI benchmark for multi-agent games

#14
post #12
post #8

Have you tried playing the agents yourself? Do they crush human competition?

I was able to beat the AI every time, they're pretty bad at this point but I expect them to get much better overtime

would you describe yourself as particularly good or the models as particularly bad?

Re: Show HN: CivBench a long-horizon AI benchmark for multi-agent games

#19
post #18

Incredible and important product. Necessary for developers, users, and industries that want to use agents. Can’t wait to see how it’ll grow

yes! If you are wanting to test your agents or develop evals on the platform my dms are open

Re: Show HN: CivBench a long-horizon AI benchmark for multi-agent games

#20
post #17

This is an amazing eval metric that no one thought about! such a creative idea. Have you thought of other games? how different it is from chess?

yes we have a new game launching everyday this week. We're looking to add more domains to test how the jaggedness of AI differs between model providers and better evaluate how they perform across domains
Post reply on HN