AI agent context compression becomes a production risk when a smaller prompt is mistaken for a preserved task state. The useful question is whether the compacted agent still makes equivalent decisions, respects the same constraints, and completes the same work. This guide provides a regression framework for testing those outcomes instead of treating token reduction as proof of quality.
Reading time: 10 minutes · About 1,950 words
TL;DR
- Treat context compaction as a lossy state migration, with tokens tracked as a resource metric.
- Compare full context twice before comparing full and compacted context. The first pair measures the model's own randomness.
- Gate releases on task success, constraint retention, side effects, cross-run reliability, and cost per successful task.
- Track state reacquisition and duplicate work. Stable completion can hide a large interaction-cost increase.
- Use decision equivalence as a diagnostic layer. End-to-end task outcomes remain the release authority.
Context compaction is a state migration
OpenAI's Compaction documentation describes two operating modes. Server-side compaction can trigger when rendered context crosses a configured threshold. A standalone /responses/compact call can instead return a canonical compacted window for the next request. In both modes, the compaction item carries prior state in fewer tokens, is encrypted, and is not intended for human interpretation.
That opacity changes the test strategy. A reviewer cannot inspect the compacted item and certify that every goal, permission, decision, identifier, and unresolved action survived. The only dependable interface is downstream behavior.
OpenAI's GPT-6 production guide points in the same direction: run representative tasks and measure task success, latency, and cost per successful task. The accompanying memory and compaction Cookbook also separates active compaction from reusable memory and reviewed artifacts. It recommends compacting at meaningful workflow boundaries and keeping cited facts in artifacts rather than only in conversation state.
The architecture should therefore assume that compaction is lossy. Durable files, event logs, approvals, and business records remain sources of truth. Compacted context is a working representation used to continue execution.
This is narrower than the system-level question in The Harness Is Part of the Benchmark. That analysis compares complete provider configurations. The protocol here isolates compaction as one versioned intervention inside a fixed harness.
Token reduction measures pressure; outcomes measure correctness
A compression ratio answers how much context disappeared. It cannot answer whether the right information survived.
Factory's vendor study illustrates the gap. Its context-compression evaluation used 36,611 messages from opted-in software-engineering sessions and tested four probe types: recall, artifact, continuation, and decision. Factory reported overall scores of 3.70 for its structured summary, 3.44 for Anthropic, and 3.35 for OpenAI under its GPT-5.2 judge. Artifact tracking was weak for every method, ranging from 2.19 to 2.45 out of 5.
Those results are useful evidence about what summaries can lose. They are also a vendor self-evaluation without a public end-to-end benchmark, raw session data, confidence intervals, or a reported human-validation rate. Probe quality is one layer of evidence, not a production release decision.
Task completion alone is also insufficient. A 2026 preprint on interaction costs under context compression found that retrieval calls increased in all six model-regime comparisons at its prespecified 5x compression point, while completion changes were not significant. In the GPT-5.5 High condition, completion moved from 80% to 85%, while retrieval calls rose from 21.0 to 63.9. The agent still finished, but it spent far more effort reacquiring state.
The correct denominator is the successful task. Count all agent calls, compressor calls, repeated reads, retries, cache effects, latency, and human repair required to produce it.
Build a decision-equivalence experiment
Start with a paired experiment that freezes everything except compaction.
Keep the model, reasoning level, system instructions, skills, tool schemas, permissions, environment snapshot, task input, step budget, and timeout constant. Version each item so a later model or prompt change cannot silently contaminate the result.
Run three arms:
| Arm | Purpose |
|---|---|
| Full context A | Production baseline |
| Full context A' | Model self-agreement and sampling-noise floor |
| Compacted context B | Compression treatment |
The A/A' comparison matters because stochastic agents can choose different valid actions from identical context. Calling every A/B difference compression damage would overstate the effect. Measure the excess decision-change rate above the A/A' floor.
A useful decision signature includes the next action, target object, authority class, and externally visible effect. Exact argument matching works for deterministic tools. Semantic normalization is needed when two calls use different syntax but produce the same authorized effect.
For example, two search queries can be equivalent while update_customer and send_refund are different authority classes. A compacted agent that changes from asking for approval to issuing a refund has crossed a much more important boundary than one that changes a search phrase.
The open-source Distil project uses this A/A plus A/B pattern in its project-reported evaluation. Its own negative result is more important than its headline metrics: a per-turn decision certificate did not transfer under aggressive compression to end-to-end task success. Decision equivalence is a sensor. The final task remains the acceptance test.
Freeze a task set that can expose state loss
Synthetic recall questions are cheap, but a production gate needs tasks with executable outcomes. Build the suite from real traces and deliberately include cases where one forgotten fact changes the correct action.
| Test family | State that must survive | Primary assertion |
|---|---|---|
| Artifact-heavy work | Files read, modified, and verified | Correct artifact and passing validation |
| Constraint-heavy work | Scope, policy, forbidden actions | Zero hard-constraint violations |
| Correction tasks | New fact that invalidates an old assumption | Latest valid fact controls the decision |
| Deferred instructions | A commitment made many turns earlier | Action occurs at the right boundary |
| Failure recovery | Failed approach and last good checkpoint | No repeated dead end or duplicate effect |
| Approval tasks | Allowed, ask, and denied authority | Same approval behavior after compaction |
Test zero, one, and repeated compaction events. Multi-cycle drift matters because a detail can survive one summary and disappear when that summary is summarized again.
Repeat each arm with multiple seeds or independent runs. Report pass@1 and a reliability metric such as the probability that all three runs succeed. Mean success can look healthy while repeated-run reliability collapses.
A 2026 counterfactual-continuation preprint shows why boundary-level diagnosis helps. Across 197 compression boundaries from 82 trajectories, 151 increased continuation length, yet only 12 added more than five steps. Mild burden was common, while severe degradation concentrated at a few boundaries. First detect regression at task level, then replay from checkpoints immediately before and after each boundary to locate the damaging compaction.
Turn hidden state into explicit assertions
Every task fixture should declare the state that matters before the run begins. A minimal state contract covers:
- Goal: the outcome and the definition of done.
- Constraints: scope limits, formats, privacy rules, and forbidden actions.
- Decisions: chosen options, rejected alternatives, and the reason a decision is settled.
- Artifacts: authoritative files, versions, source links, and validation status.
- Open work: unresolved questions, failed attempts, pending tool calls, and promised follow-ups.
- Authority: actions the agent may take, actions requiring approval, and actions it must refuse.
Convert each item into an assertion wherever possible. Compare repository hashes, database object counts, approval events, cited source IDs, test output, or a structured final-state record. LLM judges can help classify open-ended decisions, but deterministic checks should own exact identifiers and side effects.
This state contract also improves architecture. If an approval or source citation must survive every compaction, it probably belongs in a durable artifact or event log rather than only in the model's working context.
For the durable layer, AI Agent Memory Needs Source-Linked Claims describes how evidence, claims, indexes, and working context should remain separate.
Use a release scorecard, not one headline metric
Measure five groups together:
| Group | Metrics |
|---|---|
| Outcome | Task success, pass@1, repeated-run reliability, final-state diff |
| Behavior | Excess decision changes, constraint violations, approval divergence, duplicate side effects |
| Recovery | Retrieval calls, repeated reads, expand operations, repeated failed approaches |
| Efficiency | Peak context, cumulative input and output, compressor tokens, cache hits, latency |
| Economics | Total cost, cost per successful task, human repair time |
Declare the gate before running the evaluation. A practical policy might require the compacted arm to remain within a two-percentage-point non-inferiority margin on task success, produce zero additional hard-constraint or duplicate-effect failures, and reduce cost per successful task without an unacceptable rise in P95 latency or reacquisition calls.
The numbers must match the risk. A read-only research agent can tolerate a larger decision-change budget than an agent issuing payments or changing access rights. High-impact tasks need a stricter margin and deterministic authority assertions.
Do not interpret a non-significant difference as equivalence. Use paired confidence intervals or a predeclared non-inferiority test. Segment results by model, task family, context budget, and number of compaction cycles. Averages can hide a compressor that works for retrieval tasks and fails on delayed constraints.
Roll out with shadow replay and rollback
Begin in offline replay. Feed frozen traces into the full and compacted arms, inspect failures, and set the first risk budget.
Next, shadow a sample of production requests. Run the compacted branch without allowing it to create external side effects. Compare its proposed decisions with the production branch and record the A/A noise floor separately.
Then canary a small slice of low-risk work. Emit a compaction event with the session version, model, prompt, tool schema, threshold, compactor version, parent range, and artifact references. Preserve enough information to reproduce the pre-compaction state and roll back the policy.
Pause or fall back to full context when any hard assertion fails, the upper confidence bound exceeds the decision-change budget, repeated-run reliability drops, or cost per successful task rises. Token savings alone never override those gates.
FAQ
What is context compaction in an AI agent?
It replaces part of a long interaction history with a smaller representation that carries state into later turns. It is working-state management, not the same thing as durable memory or an authoritative event log.
When should an agent compact context?
Use a threshold or a meaningful workflow boundary before the active context becomes operationally expensive or reaches the model limit. Choose the threshold from task-level regression results, because the best setting changes with the model, tools, and task.
Does context compression always reduce cost?
Savings depend on the complete workflow. A smaller prompt can still increase compressor calls, cache misses, repeated reads, recovery calls, retries, or failures. Measure cumulative cost per successful task.
What information is most dangerous to lose?
Hard constraints, approvals, corrections, artifact identifiers, settled decisions, failed approaches, and pending side effects. These items change what the agent is allowed or expected to do next.
Is decision equivalence enough?
Decision equivalence is an early diagnostic for behavioral drift. End-to-end outcomes, repeated-run reliability, constraint checks, side-effect checks, and total-cost accounting provide the production release gate.
The next action
Take ten recent long-running traces, freeze their environments, and write the state assertions before changing the compressor. Run full context twice and the compacted policy once for each trace. The first useful result is not a compression ratio. It is a list of the exact decisions, constraints, and recovery costs that changed.
References
- OpenAI API: Compaction
- OpenAI: A model guide for the GPT-6 family
- OpenAI API deployment checklist
- OpenAI Cookbook: Building Reliable Agents with Memory and Compaction
- An Empirical Study of Harness Design for Coding Agents
- What Does Context Compression Cost an Agent?
- Adapting Context Compression for Long-Horizon Agents with Counterfactual Continuations
- Factory: Evaluating Context Compression for AI Agents
- Distil project evaluation