OpenGameEval: Eval Framework to Benchmark Agentic AI Assistants
1–2 of 2 posts
Re: OpenGameEval: Eval Framework to Benchmark Agentic AI Assistants
#2OpenGameEval offers a unique testing ground to evaluate core model capabilities related to agenetic reasoning and long-horizon task solving.