Administrator
Published on 2026-10-10 / 4 Visits
0
0

ThinkingBox Agent Benchmark Audit: What 20 Runs Reveal Beyond Pass@1

The ThinkingBox agent benchmark repeats 507 stateful business tasks 20 times and grades the database state left behind. Its important result is larger than a leaderboard: discovering one successful trajectory and delivering dependable execution are different capabilities. Yet an observed 20-for-20 result is still a sample, not a production SLA. This audit separates the metrics, verifies the released artifacts, and marks the boundary between reproducible benchmark design and independently reproduced historical results.

What ThinkingBox actually releases

ThinkingBox is the runtime. It orchestrates an LLM agent, a simulated user, and isolated MCP-compatible tool sessions, then passes terminal state and recorded side effects to executable judges. ThinkingBox-Bench v1.0 is the frozen evaluation release.

The release manifest contains exactly 507 task instances across five domains:

Domain Tasks
Retail and e-commerce 98
Travel and hospitality 104
Auto insurance 100
Neobank support 104
Consulting IT and HR support 101
Total 507

Each task defines an initial backend state, a user goal, a domain policy, available tools, private facts held by the simulated user, and hidden executable checks. Every attempt starts from a clean state in an isolated session. The evaluator compares the resulting database and side effects with the required outcome. All checks are conjunctive, so one wrong field, missing effect, or unauthorized extra effect fails the task.

All 507 tasks inspect backend state. Thirty also add binary response rubrics for requirements such as disclosure, confidentiality, or consistency between the action and the final message. The remaining 477 deliberately prioritize deterministic backend evidence over a semantic judgment of the response.

That design answers a specific question: did the agent leave the business system in the required state? It does not claim to measure every property of a good customer interaction.

Four metrics answer four different questions

Several reports compress ThinkingBox into a comparison between pass@1 and 20-for-20. The current paper v4 contains four distinct quantities.

Metric Question answered Calculation boundary
pass@1 How often does one complete attempt succeed? Micro-average of 507 tasks times 20 trials
pass@20 Can 20 attempts discover at least one successful trajectory? At least one observed success per task
plug-in pass^20 How much repeated-success probability is estimated from per-task rates? Average of (successes / 20)^20 across tasks
observed 20/20 Which tasks passed every recorded attempt? Literal count with 20 successes out of 20

The notation matters. In v4, pass^20 is a biased plug-in estimator chosen because the unbiased estimator becomes zero whenever a task records fewer than 20 successes. Observed 20/20 is a harsher empirical count with no smoothing.

GPT-5.4 shows why these values cannot be substituted for one another:

  • pass@1: 65.36%
  • pass@20: 91.12%, meaning 462 of 507 tasks succeeded at least once
  • plug-in pass^20: 30.62%
  • observed 20/20: 128 of 507 tasks, or 25.25%

An earlier Microsoft explanation labels 25.25% as pass^20, while paper v4 reports 30.62% under the plug-in definition and lists the 128 literal all-success tasks separately. Any comparison needs a version date and an explicit metric definition.

One more trap is mathematically important. Do not take the aggregate pass@1 and raise it to the twentieth power. For GPT-5.4, 0.6536^20 is about 0.020%, far below the paper's 30.62% task-level average. The gap reveals extreme task heterogeneity. Some workflows are nearly always solved; others are almost never solved. Averaging first destroys that structure.

For the broader framework around configuration identity, grader validity, repeated trials, and deployment evidence, see the complete AI agent evaluation framework.

Capability breadth and dependable repetition rank models differently

The v4 results contain a useful rank inversion.

Model pass@1 pass@20 plug-in pass^20 Observed 20/20 tasks
Claude Opus 5 66.50% 79.09% 47.53% 241
GPT-5.4 65.36% 91.12% 30.62% 128
GPT-6 Astra 58.31% 71.01% 46.89% 231
Kimi-K3 57.37% 93.89% 17.60% 68

Kimi-K3 finds a successful route on the widest share of tasks, but reproduces success much less consistently. GPT-6 Astra ranks fifth on pass@1 in the paper, yet second on plug-in pass^20. Model selection changes when the product needs one recoverable success, a strong first attempt, or reliable repetition.

The October 3 Hugging Face article adds a later official result for Claude Opus 5.5: 67.16% pass@1 and 241 observed 20/20 tasks. That model is absent from the October 1 paper v4 table, so it should be treated as a post-paper addition rather than silently merged into the paper's experiment.

A valid tool call is still only process evidence

ThinkingBox's strongest evidence comes from comparing surface success with terminal state.

In a common-set retrospective covering 121,680 valid trials from 12 models, 79,853 trials failed the executable checks. Among those failures, 67.24% still ended cleanly, invoked a state-changing tool, and showed no final tool error. State inspection found wrong field values in 77.61%, unintended extra effects in 43.30%, and missing required effects in 25.36%. Those categories overlap.

The paper also assigns each failed trace one deterministic diagnostic signature. Across models, the unweighted average distribution is:

  • tool usage and failed recovery: 79.9%
  • wrong state update: 10.3%
  • incomplete user resolution: 7.0%
  • no state-changing action: 2.9%

These are observable labels, not unique causal explanations. Still, they change where engineering effort should go. A model that knows the policy may still mishandle a failed precondition, continue after an empty lookup, update the wrong object, or stop before the requested mutation.

The agent's final statement is a claim. A valid call is another claim about progress. The durable postcondition is the acceptance evidence.

What 20-for-20 does and does not prove

Twenty clean successes are a valuable screening signal. They do not prove a true success probability of 100%.

Under a simplified independent Bernoulli model, observing zero failures in 20 trials gives a one-sided 95% lower bound of approximately 0.05^(1/20) = 86.1% for the underlying success probability. Production failures are often correlated, so even that interval may travel poorly across time, traffic, providers, and changed user behavior.

ThinkingBox resets every attempt to the same initial state. This isolates repeated execution and prevents earlier runs from contaminating later ones. It does not test twenty consecutive production events that accumulate state, share rate limits, encounter schema migrations, or experience a provider incident.

The benchmark has other explicit limits:

  • tasks are synthetic reconstructions from a non-public source collection, not a random sample of enterprise work;
  • each retained task has one golden terminal state, so workflows with several defensible resolutions are excluded;
  • the simulated user is cooperative, has a fixed objective, and permits at most ten follow-up turns;
  • 477 tasks can pass even if the agent changes the backend correctly but describes the result incorrectly;
  • the same GPT-5.4-mini setup serves as the simulated user and the judge for the 30 response-rubric tasks;
  • isolation is a documented design property, while the public evidence does not include an independent cross-trial contamination study.

These limits narrow the claim. ThinkingBox measures reproducible completion of defined, stateful, policy-conditioned workflows under its recorded harness. It does not certify general enterprise reliability.

Re-runnable is not the same as independently reproduced

The release is unusually inspectable. It publishes the runtime, task definitions, MCP servers, a frozen 507-item manifest, dataset views, aggregation code, and instructions for generating one JSONL row per trial. The implementation confirms that pass^k is calculated as the average task-level plug-in estimate.

What is not currently present in the public release is equally important: the complete historical JSONL for all reported model runs. The paper includes a small set of representative full traces, but the repositories and dataset do not expose the full 18-model trial corpus with conversations, tool outputs, token records, and per-trial verdicts.

An external team can run a new experiment. It cannot reconstruct every published aggregate, bootstrap interval, failure label, or possible batch correlation from the released historical records alone. Closed model updates make exact historical reproduction harder over time.

This is the same evidence distinction that matters when using public checks for program repair. The boundary between published checks and independent evaluation should remain visible.

Turn the benchmark into a production reliability contract

ThinkingBox is most useful as a design pattern, not as a universal score threshold.

  1. Freeze representative workflows, initial states, policies, tools, graders, model routes, and budgets.
  2. Grade terminal state, missing effects, collateral effects, permissions, and response fidelity separately.
  3. Report pass@1, allowed-retry recovery, task-level consistency, latency, cost, and human intervention.
  4. Preserve the raw trial record so every aggregate can be recomputed and every failure can be inspected.
  5. Test correlated failure modes such as provider outages, stale credentials, schema changes, concurrency, and shared rate limits.
  6. Choose repetition counts and acceptance thresholds from failure cost rather than copying 20 by default.
  7. Promote only the stable subset into shadow traffic, then canary it against real outcomes and rollback signals.

Teams that need an implementation path can use the guide for turning product state into a dynamic release gate.

The deeper lesson is not that every agent needs exactly twenty trials. It is that reliability must be attached to a defined task, a frozen system configuration, an external verifier, and a declared time horizon. One success demonstrates capability. Repeated, inspectable success earns a narrower form of trust.

Frequently asked questions

What is the difference between ThinkingBox and ThinkingBox-Bench?

ThinkingBox is the reusable runtime and evaluation harness. ThinkingBox-Bench is the frozen set of 507 executable business tasks, tool servers, policies, and checks used for the published evaluation.

Does stateful mean long-term agent memory?

Here it primarily means that tools mutate persistent business state during a task. The evaluator inspects the resulting database and side effects. The benchmark does not establish long-term memory reliability across unrelated production tasks.

Does passing 20 out of 20 prove 100% reliability?

No. It proves zero observed failures in that specific 20-trial sample. The true rate remains uncertain, and production conditions can introduce correlated failures absent from the benchmark.

Why can correct tool calls still fail?

A call can target the wrong record, use the wrong field value, omit a required follow-up action, create an extra effect, or fail to recover from a precondition. Tool validity and business completion are different evidence levels.

Did ThinkingBox invent terminal-state grading and pass^k?

No. Earlier work such as tau-bench used terminal database state and repeated-success metrics. ThinkingBox's contribution is the combined MCP-compatible runtime, five-domain task set, explicit side-effect checks, frozen release, and large repeated-trial campaign.

References


Comment