Administrator
Published on 2026-09-22 / 20 Visits
0
0

AI Agent Memory Needs Source-Linked Claims

AI agent memory becomes trustworthy when every actionable claim keeps its source, validity, dependencies, and invalidation conditions. A larger context window can expose more text, but it cannot determine which statement is current or which downstream decision should change after a source is corrected. This guide turns institutional memory into a source-linked claim system that teams can test, audit, and repair.

Reading time: 8 minutes · About 1,600 words

TL;DR

  • Store claims with source spans, time boundaries, authority, dependencies, and review state.
  • Keep raw evidence, derived claims, retrieval indexes, and agent context as separate layers.
  • Treat retrieval as candidate selection. Validate a claim before using it for a consequential action.
  • Propagate corrections through dependency links instead of silently overwriting old memory.
  • Measure memory quality through provenance coverage, conflict handling, freshness, and downstream action accuracy.

V7 shows the useful boundary of a context graph

OpenAI's V7 case study, published on September 21, 2026, describes a Context Graph that connects entities, relationships, and cited evidence. V7 Go extracts information from company repositories, keeps recent exchanges in active model context, and retrieves older material from the graph when needed. The V7 product page describes the same layer as firm memory spanning reports, data rooms, memos, emails, entities, relationships, and source evidence.

This architecture addresses a real limitation. An organization owns years of history, while an agent receives a bounded working set for one task. Repeating document search on every request wastes time and makes important relationships easy to miss. A persistent graph can preserve the connections that retrieval needs.

The important engineering distinction is that a graph entry is still a claim. A graph edge can be stale, derived from a weak source, contradicted by a newer filing, or created by an extraction error. Connectivity improves retrieval. Provenance and lifecycle controls determine whether the retrieved statement deserves trust.

This also separates the topic from our earlier article on latent workspace versus external memory. That article draws the state boundary between model computation and durable storage. The next problem is narrower: what must durable storage preserve so that another agent can safely reuse it?

Make the claim, not the document chunk, the memory unit

Traditional retrieval systems often store chunks plus embeddings. The chunk answers where a passage lives. An institutional memory system also needs to answer what assertion was derived from that passage and when the assertion applies.

A minimal claim record can look like this:

{
  "claim_id": "claim:fund-17:nav:2026-q2",
  "subject": "fund-17",
  "predicate": "reported_nav",
  "value": {"amount": 184000000, "currency": "USD"},
  "source_uri": "drive://reports/fund-17-q2-2026.pdf",
  "source_span": {"page": 12, "table": "Portfolio Summary"},
  "observed_at": "2026-07-19T09:14:00Z",
  "valid_from": "2026-06-30",
  "valid_to": null,
  "authority": "fund_manager_report",
  "derived_from": [],
  "status": "reviewed",
  "revalidate_on": ["source_changed", "newer_period_arrived"]
}

The exact schema will vary. The invariants matter more than the field names:

  1. The claim has stable identity.
  2. The original evidence remains addressable at passage level.
  3. Observation time and business validity are separate.
  4. Derived claims name their dependencies.
  5. Status can represent proposed, reviewed, disputed, superseded, or invalid.
  6. Revalidation has explicit triggers.

This design follows a simple principle: a system becomes verifiable when errors have somewhere visible to attach. A free-form summary hides the failing sentence inside prose. A claim record gives the correction, reviewer, and dependency graph a stable target.

The W3C PROV-O recommendation offers a useful general vocabulary for entities, activities, and agents. Agent memory needs additional domain fields for authority, validity, conflict, and action risk, but it can preserve the same separation between evidence, transformation, and responsible actor.

Separate evidence, claims, indexes, and working context

One storage layer should not perform four different jobs. A safer architecture separates them:

Layer Purpose Typical failure
Raw evidence Preserve the received source or immutable snapshot Source disappears or changes silently
Claim store Represent bounded assertions and their lifecycle Unsupported or stale assertion remains active
Retrieval index Find relevant evidence and claims efficiently Ranking omits a decisive item
Working context Give the model a small task-specific view Too much low-signal material crowds out constraints

The raw layer is the audit base. The claim store is the reusable institutional memory. The index is a replaceable performance structure. Working context is a temporary projection for the current decision.

This separation prevents a common failure: an embedding index becomes the de facto source of truth. Indexes are derived artifacts. They may lag behind deletes, corrections, access-control changes, or entity merges. Rebuilding an index should change retrieval performance while leaving authoritative evidence and claim history intact.

The same distinction appears in verifiable analytics. Lineage routes a reader to evidence; validation checks whether the transformation and interpretation were sound. A complete path to the wrong claim remains a traceable error.

Use different contracts for reading and writing memory

Memory reads optimize for relevance under a token budget. Memory writes change what future agents may believe. They require different controls.

For reads, retrieve a bounded subgraph and include:

  • the candidate claim and its current status;
  • the supporting source span;
  • newer, conflicting, or superseding claims;
  • the reason this item matched the task;
  • the access policy that permits its use.

For writes, require:

  • a stable source or an explicit observation event;
  • an extraction or derivation method and version;
  • entity-resolution confidence;
  • a rule for conflicts with existing claims;
  • a reviewer or automated acceptance test for high-impact fields;
  • a revalidation trigger.

An agent conversation should rarely become institutional truth by default. A conversation can create a proposed memory. Promotion into a reviewed claim requires evidence and an owner. This preserves useful learning while preventing plausible model output from entering the firm's history as fact.

Research on evidence tracing and execution provenance in LLM agents frames the same need at system level: documents, observations, memory items, intermediate claims, actions, and final answers form a dependency network. That network enables selective invalidation and shows which evidence influenced an action.

Corrections must propagate

Institutional memory fails quietly when a source changes but derived summaries remain active. A source link alone supports inspection. A dependency link supports repair.

Suppose a quarterly report corrects a portfolio company's revenue. The system should perform four steps:

  1. Create a new source snapshot and retain the earlier one.
  2. Supersede the old revenue claim with an explicit reason.
  3. Mark derived claims, memos, and cached answers for revalidation.
  4. Recompute only the affected projections and notify owners of consequential actions.

Avoid deleting the old claim as if it never existed. Historical decisions may have been reasonable under the evidence available at the time. Temporal history supports both audit and learning.

Conflicts also need a first-class state. Two valid sources may disagree because they use different dates, accounting definitions, or authority levels. Retrieval should expose that disagreement. Silent winner selection turns a data-quality question into hidden model policy.

An implementation gate for institutional memory

Before letting an agent use memory for decisions, test one end-to-end claim path:

  1. Ingest: Can the system freeze or version the original source?
  2. Extract: Can a reviewer jump from the claim to the exact supporting span?
  3. Resolve: Can the system distinguish similar entities and explain merges?
  4. Conflict: Does a contradictory source create a visible contested state?
  5. Retrieve: Does the task context include status, source, and freshness?
  6. Act: Does action policy require stronger evidence as impact rises?
  7. Invalidate: Does a source correction reach every dependent claim and output?
  8. Audit: Can the team reconstruct what the agent knew when it acted?

Track at least four metrics: source-link coverage, stale-claim rate, unresolved-conflict rate, and consequential actions supported by reviewed claims. Recall and latency remain useful, but they measure access rather than trust.

Start with one narrow workflow such as investment memo updates, support-policy answers, or vendor-risk reviews. A bounded workflow exposes the lifecycle rules quickly. Scaling storage before proving invalidation creates a larger memory with a larger silent-error surface.

FAQ

Is a context graph the same as agent memory?

A context graph can implement part of agent memory by linking entities, claims, relationships, and evidence. A production memory system also needs write policy, validity, conflict handling, access control, revalidation, and audit.

Does a larger context window remove the need for long-term memory?

A larger window increases the material available to one inference. Long-term memory preserves selected state across tasks and time. Source binding and invalidation remain necessary at any context size.

Is provenance enough to prevent hallucinations?

Provenance makes a claim inspectable. Validation still has to check extraction, interpretation, freshness, and fit for the current decision. Provenance is a route to evidence, not a correctness certificate.

Should an agent save every conversation?

Keep raw logs according to retention policy, then promote only reusable, supported claims into institutional memory. High-volume transcript storage and trusted memory serve different purposes.

What happens when two sources disagree?

Preserve both claims, attach dates and authority, mark the conflict, and apply a domain-specific resolution policy. High-impact unresolved conflicts should block automated action.

Build the repair path before the memory grows

Institutional memory should help an agent inherit organizational judgment rather than accumulate text. Source-linked claims provide the missing unit: small enough to validate, stable enough to reference, and connected enough to repair when evidence changes.

Choose one consequential claim in an existing workflow. Trace its source, validity, dependencies, and invalidation path. If any link is implicit, that link is the next memory feature to build.

References


Comment