MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks 2025
1–2 of 2 posts
Re: MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks 2025
#2How do you handle partial success in the benchmark? If an agent picks the wrong tool mid-chain but fixes it with backtracking, does it still count as a full success?