MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks 2025

3 pointsposted 5 hours ago
by janandonly

1 Comments

dima853

5 hours ago

How do you handle partial success in the benchmark? If an agent picks the wrong tool mid-chain but fixes it with backtracking, does it still count as a full success?