An AI research agent can find the best configuration and still misunderstand most of the system around it. WhatWorkedBench makes that gap measurable by scoring both configuration choice and predictions for unobserved interventions. The practical lesson is simple: evaluate exploration and reconstruction separately before reusing an agent's conclusions.
Reading time: 8 minutes · Length: about 1,650 words
TL;DR
- WhatWorkedBench asks agents to buy a limited number of experiments and predict the complete configuration response surface.
- On 22 four-factor sources at an eight-measurement budget, pair-effect ridge found the exact optimum 15 times but passed strict reconstruction only three times.
- Thirteen sources combined an exact best choice with at least one conditional-effect error above 10% of the true utility range.
- A production research agent therefore needs two scorecards: search quality and system reconstruction.
- Code-derived constraints help only when they are enforced in the final estimator or artifact.
The winner is a weak proof of understanding
Suppose an agent tests ten training recipes and returns the strongest one. The result is operationally useful. It may save compute or improve a model. Yet it leaves a different question unanswered: did the agent learn why the recipe worked, or did its search happen to land on a high point?
That distinction matters as soon as a team asks the agent to transfer a lesson. A rule such as “enable component A” may help in one background and hurt in another. If the agent only knows the winning combination, it cannot reliably predict the next intervention, explain an interaction, or recognize when the original result has stopped applying.
WhatWorkedBench turns this distinction into an executable benchmark. An agent sees workflow code, receives two free measurements, purchases a limited number of additional results, and submits predicted scores for every legal configuration. Exhaustive native CPU runs provide the hidden reference surface.
This is a stronger contract than “beat the baseline.” It asks what the agent learned from the experiments it chose.
What the benchmark actually measures
The public release covers 36 task conditions, 30 sources, eight workflow families, and 1,248 indexed configuration records. The families include classification, regression, clustering, forecasting, image restoration, retrieval, beat detection, and graph link prediction. The repository also records 4,206 numerical-control executions and provides native replay and audit scripts.
Four-factor tasks have 16 configurations. Six-factor tasks have 64. For each task, the agent gets the all-off and all-on configurations as free anchors, then spends a budget on additional measurements. Its final artifact predicts the complete table.
The evaluator scores several different properties:
| Target | Question |
|---|---|
| Exact optimum and selection regret | Did the agent choose a good final configuration? |
| Conditional-effect error | Can it predict the effect of changing one component in every background? |
| Pair-interaction error | Does it understand when two changes depend on each other? |
| Grid error | Can it reconstruct the complete response surface? |
| Delivery | Did it produce a valid, complete artifact? |
The conditional-effect metric is the key addition. It evaluates a component change while holding all other options fixed. This exposes reversals that a single average effect hides. The paper reports meaningful sign reversals across much of the catalog: a switch can improve one background and degrade another.
The central result: optimization and understanding separate
At a budget of eight purchased measurements on 22 four-factor sources, pair-effect design with ridge reconstruction achieved a family-macro effect recovery of 0.612. It selected the exact optimum on 15 sources.
That sounds strong until the reconstruction test is added. Only three sources passed strict reconstruction, which requires every conditional-effect error to stay within 10% of the true utility range. Thirteen sources combined an exact optimum with at least one error above that threshold.
The numbers support a precise conclusion. Selecting the winner and reconstructing the system are complementary targets. Neither can substitute for the other.
The comparison with an effect-variance Gaussian process reinforces the point. At the same budget it reached 0.701 recovery and selected 16 exact optima, yet passed strict reconstruction on only one source. A higher average recovery score, more correct winners, and fewer all-effects passes can coexist because the metrics inspect different failure shapes.
The correct response is not to crown one metric. It is to keep the scorecard multidimensional.
Separate evidence acquisition from inference
A research agent makes at least two algorithmic decisions:
- Which experiments should it run?
- How should it infer unobserved results from those measurements?
End-to-end scores mix those decisions. WhatWorkedBench also fits shared estimators to fixed observation sets. This lets the authors ask whether an agent bought informative evidence and whether its chosen reconstruction method used that evidence well.
The results resist a simple model ranking. In a prospective typed study across 12 sources, Flash delivered 12 of 12 direct tables and exceeded a Gaussian process fitted to the same observations by 0.147 recovery on average. Pro delivered 11 of 12; its all-attempt difference was -0.001, while its delivered-only difference was +0.063.
These figures show why delivery belongs in the evaluation contract. A strong prediction that never arrives has zero operational value. They also show why one aggregate score can mislead: acquisition, inference, and artifact completion can move independently.
Program structure is evidence, but only when enforced
Some configurations are equivalent by construction. A parameter may become inactive when its parent feature is disabled. An agent that reads the code can identify this rule without buying duplicate measurements.
Recognition alone is insufficient. In a separate diagnostic, agents reported all registered code relations, yet their final ridge-helper tables still violated native equivalences. Applying a class-consistent projection after the fact raised recovery from 0.338 to 0.507.
In fixed-observation runs, fresh agents passed all six rules into named estimators. All 96 equivalent configuration pairs then agreed. On those observations, verified rules raised D2 recovery from 0.319 to 0.601 and GP recovery from 0.509 to 0.666.
This is a general engineering lesson: a rule in the agent's reasoning trace is only an observation. A rule encoded in the estimator or final-artifact validator becomes a system property.
A two-scorecard contract for research agents
Teams building research agents can adapt the benchmark's separation without copying the entire benchmark.
Scorecard 1: exploration utility
Measure whether the agent uses its budget well:
- best observed or selected utility;
- selection regret against a reference when one exists;
- cost, wall time, and failed experiments;
- diversity and coverage of tested configurations.
This scorecard answers whether the run produced a useful candidate.
Scorecard 2: reconstruction quality
Hold back configurations and ask the final model to predict them. Measure:
- error on unobserved outcomes;
- conditional effects across different backgrounds;
- interaction errors;
- calibration or uncertainty on held-out regions;
- consistency with known code or domain constraints.
This scorecard answers whether the run produced reusable knowledge.
The distinction should persist in reporting. A result can be labeled “best configuration found” without being labeled “mechanism understood.” This prevents a successful search episode from silently acquiring more epistemic status than its evidence supports.
A minimal implementation protocol
An internal evaluation can start with six steps:
- Freeze the task. Fix code, data split, seed policy, metric, and legal interventions.
- Reserve a reference set. Keep several configurations or intervention contrasts hidden from the agent.
- Cap the measurement budget. Record every query, including failed and repeated attempts.
- Require a complete artifact. Ask for a table or predictive function, not only a recommended winner.
- Apply deterministic constraints. Enforce equivalence classes, inactive parameters, ranges, and schema checks outside the model.
- Report both scorecards. Publish search performance, held-out reconstruction, delivery status, and the exact evidence boundary.
The WhatWorkedBench repository demonstrates this separation with a trusted evaluator process, public agent inputs, hidden references, reproducible native runs, and machine-readable audits. Its exact protocol is specialized, but the boundary design is broadly reusable.
What the evidence does not establish
The benchmark uses deterministic CPU workflows, binary component switches, and small exhaustive configuration spaces. That makes rigorous reference construction possible. Open-ended scientific research introduces changing hypotheses, measurement noise, ambiguous objectives, and interventions that cannot be exhaustively enumerated.
The paper therefore does not prove that a response-surface test is sufficient for science. It proves a narrower and valuable point: even in a controlled environment, finding the best observed configuration cannot certify broad experimental understanding.
That is enough to change how research agents are evaluated today.
FAQ
What is experimental understanding in an AI research agent?
In this benchmark, it is the ability to predict how documented component changes affect a workflow after a limited set of experiments. It is measured across the response surface, not inferred from the final winner alone.
Why is choosing the best configuration insufficient?
Several different and partly incorrect models can rank the same configuration first. They may disagree everywhere else, especially when component effects reverse across backgrounds.
How can a team test understanding without exhaustive experiments?
Reserve a hidden set of configurations or intervention contrasts, require predictions before revealing outcomes, and score both outcome error and conditional-effect error. Known structural constraints can provide additional deterministic checks.
Does a Gaussian process always outperform the agent?
No. The paper reports cohorts where direct agent tables beat a GP fitted to the same observations, as well as conditions where shared estimators improve reconstruction. The useful design is the controlled comparison, not a universal winner.
Can this protocol evaluate real scientific agents?
Parts of it can. Budget logging, held-out interventions, complete artifacts, constraint checks, and separate exploration/reconstruction metrics transfer well. Exhaustive reference surfaces usually do not.
References
- WhatWorkedBench paper, arXiv v2
- WhatWorkedBench public repository
- Evaluation protocol
- Claude-Shaped Science: Build the Research Harness Around Verifiability
- GEPA optimize_anything: The Evaluator Is the Interface
Next action: take one research-agent benchmark you already use and add a held-out reconstruction score beside its best-result score. The difference between the two is the beginning of an honest capability map.