Coding-agent benchmarks often start from repositories and issue descriptions. Failure Map starts from a smaller unit: a behavior contract, an implementation that violates it, an unsuccessful repair, and executable boundary checks.
Release 2026.09.5 reports 20,168 open Python debugging tasks across 254 software-failure categories. The number is interesting. The methodology is more useful because it makes several distinctions that benchmark builders routinely blur: mechanism versus variant, public check versus hidden evaluation, and recorded execution versus general correctness.
What an open task contains
The public gzip export includes a prompt, broken source, a hard negative representing an unsuccessful fix, a Python environment contract, evaluation grouping, and recorded-boundary reward metadata. Reference solutions are null in the open tier.
The first published examples capture familiar boundary failures:
- deduplicate events by identity without collapsing distinct events with equal values;
- treat expiration intervals as inclusive at insertion and exclusive at expiration;
- paginate using the complete ordered key, including a unique identifier.
These are small programs, but they test semantic distinctions that frequently survive ordinary happy-path testing.
Failure Map says its broader archive contains 100,840 variants covering 20,168 distinct mechanisms. Five numbered cases in a family share a contract. The project also reports executing 302,520 implementations in separate Python processes and recording stdout, stderr, exit status, elapsed time, observations, and source hashes.
Those are dataset-provider claims documented in the methodology page. They make the artifacts inspectable; they do not establish model quality.
Public checks are not a hidden benchmark
The project's most important warning is explicit: recorded checks are not an independent hidden benchmark.
Once a model sees the checks, it can optimize directly for them. Passing proves behavior on those fixtures. It does not prove the repair handles broader inputs, integration conditions, or unseen variants.
This creates four evidence levels:
- Executable: the task runs in the declared environment.
- Fixture-correct: the patch passes published checks.
- Held-out robust: it passes unseen tests isolated from the model.
- Production-relevant: it survives repository integration, nonfunctional requirements, and realistic workflows.
Reporting a level-two result as level four is benchmark inflation.
Keep families together or leak the answer
Random row splitting is unsafe when variants share a contract, solution pattern, or evaluation group. A model can encounter one member during training and a near-duplicate during testing, producing a strong score without demonstrating transfer to a new failure mechanism.
Failure Map instructs users to keep related families and shared evaluation_group values together. A sound split should therefore be group-aware:
mechanism family -> exactly one of train, validation, or test
shared evaluation_group -> exactly one split
hidden fixtures -> test infrastructure only
reference repair -> inaccessible to the model
Report both mechanism coverage and variant counts. Twenty thousand rows can still represent far fewer independent ideas.
A six-step evaluation protocol
1. Freeze the task release
Record dataset version, export checksum, Python version, dependencies, and task IDs. A moving catalog makes comparisons impossible.
2. Build group-aware splits
Split by mechanism family and evaluation group before any model sees the data. Audit prompts and source for near-duplicates across splits.
3. Isolate execution
Run each candidate patch in a disposable process or container with CPU, memory, time, filesystem, and network limits. Treat model-generated code as untrusted.
4. Separate visible and hidden checks
Visible tests support development. Hidden tests support evaluation. Include property-based, metamorphic, and integration cases where the contract permits them.
5. Score more than pass rate
Track compile or parse success, visible-test pass rate, hidden-test pass rate, regression rate, patch size, attempts, wall time, and compute cost. A patch that deletes checks should fail integrity validation before scoring.
6. Preserve the evidence trail
Store the original task, generated patch, stdout, stderr, exit code, resource use, test identities, model version, prompt, and harness version. Aggregate scores without artifacts are hard to audit.
Use failure mechanisms to grow the suite
A useful evaluation set evolves from real misses. When an agent patch fails code review or production validation, reduce the incident to its smallest behavior contract, add a public training example if appropriate, and keep a distinct hidden variant for regression testing.
This creates a flywheel:
production failure -> minimized contract -> isolated task -> hidden regression -> release gate
The goal is not to maximize task count. It is to increase coverage of failure mechanisms that matter to the target workflow.
What Failure Map does and does not establish
Failure Map provides an inspectable source of controlled program-repair tasks and unusually clear split warnings. Its open tier intentionally omits passing repairs. The recorded checks establish only the stated contract and fixtures. The provider makes no model-performance improvement claim.
That restraint is a feature. It allows teams to use the archive as raw material for a defensible evaluation instead of mistaking a downloadable dataset for a finished benchmark.
FAQ
Is Failure Map a ready-made hidden benchmark?
No. Its recorded checks are public and explicitly described as non-independent. You must create and protect separate held-out tests.
Are all 20,168 tasks independent?
The project distinguishes mechanisms, variants, families, and shared evaluation groups. Evaluation should report these separately and split by group.
Why include an unsuccessful repair?
It creates a hard negative that catches plausible local fixes which still violate at least one boundary condition.
Can the tasks be executed safely on a developer laptop?
The published sources use the Python standard library, but generated code remains untrusted. Use resource-limited isolation and disable unnecessary network and filesystem access.