AI agent benchmark scores mix models, harnesses, tasks, tests, and environments. Use this framework to build a reproducible coding-agent eval.