Astra scored 62.7% or 99.9% on ARC-AGI-3 under different harnesses. Here is what changed, what the numbers mean, and what remains unproven.
AI agent benchmark scores mix models, harnesses, tasks, tests, and environments. Use this framework to build a reproducible coding-agent eval.