Administrator
Published on 2026-09-25 / 19 Visits
0
0

Astra for Law Needs a Matter-Level Audit Trail

Better answers are the starting point. Legal AI also needs a matter-level audit trail. Astra for Law improves legal research retrieval, Cooley GO Public orchestrates an IPO workflow, and Harvey brings matter context into drafting. The remaining challenge is proving which sources supported each claim, who reviewed it, which version was approved, and how an error was corrected.

9-minute read · approximately 1,900 words

TL;DR

  • Astra for Law passed an overall correctness check on 54.0% of 200 private validation questions, versus 38.7% for GPT-6 Astra with web search. That is a 15.3-point absolute gain, not proof of production reliability.
  • Astra for Law, GO Public, and Harvey expose three layers of a legal AI system: authority retrieval, controlled workflow, and matter-context drafting.
  • None of the public case studies documents a complete matter-level audit trail from authorization through final delivery and correction.
  • A legal AI evidence contract should connect each claim to an exact source passage, authority status, reviewer decision, approved document version, and reopening condition.
  • Citations are verification interfaces. Professional responsibility remains with the lawyer and the firm.

Evidence reviewed on September 25, 2026. Product claims, independent benchmark evidence, professional guidance, and the proposed controls below are labeled separately. This is systems analysis, not legal advice.

What Astra for Law's 54% actually measures

OpenAI describes Astra for Law as a configuration of GPT-6 Astra, a legal search index, specialized instructions, tools, and settings. Its index searches U.S. case law, statutes, regulations, court rules, and administrative decisions across more than 230 million URLs.

OpenAI tested that complete configuration on 200 questions from the private validation split of Vals AI's Legal Research Bench. At the highest reasoning effort, it passed an overall correctness check on 54.0% of questions, compared with 38.7% for GPT-6 Astra using web search alone.

The result supports a narrow conclusion: a legal index plus a legal workflow materially outperformed a web-search-only baseline on this validation set. The absolute gain was 15.3 percentage points. The relative gain was about 39.5%, which OpenAI rounded to 40%.

It also sets a boundary. A 54.0% pass rate means 46% did not pass that check. OpenAI did not publish confidence intervals, run counts, cost, latency, the complete configuration, or the size of the audited passage subset.

Two other figures describe retrieval rather than overall correctness:

  • On case-law-focused questions, Astra for Law found 24% more reference cases.
  • On an audited target-passage subset, it retrieved up to 54% more relevant passages from the correct opinions.

That second 54% is a maximum relative retrieval gain. It is not the 54.0% correctness rate.

The Vals benchmark documentation adds another important constraint. The benchmark contains 413 questions: 5 public, 200 private validation, and 208 test questions. The public leaderboard uses the test split. OpenAI used the private validation split, so the scores cannot be compared directly with the public ranking.

This is why benchmark evidence belongs inside an audit trail. A number without its dataset split, system configuration, evaluator, date, and failure rule is easy to overclaim.

Three cases reveal three system layers

The three September releases are more useful when read as parts of one stack.

Layer 1: authority retrieval

Astra for Law searches public legal authority and returns passages a lawyer can inspect. OpenAI says its CourtListener integration includes a collection covering more than 99.9% of published U.S. precedential case law. The Free Law Project coverage statement applies specifically to published precedential opinions. It does not cover every unpublished opinion, docket filing, contract, client fact, or commercial database.

Corpus coverage is an input property. It does not prove that a retrieved case remains good law, matches the jurisdiction, supports the proposition, or applies to the client's facts.

Layer 2: controlled workflow

Cooley GO Public begins with client information, public sources, and curated precedents to produce a tailored IPO starting point. Its agentic harness specifies which steps can run automatically and where lawyers must review or validate the work.

Cooley's own announcement says lawyers remain involved in legal analysis and the final work product. It also says an initial Form S-1 draft can be produced in minutes rather than days. That timing claim has no published sample size, baseline protocol, revision volume, or final filing-cycle measure, so it should be treated as a customer claim rather than an independently verified outcome.

Layer 3: matter-context drafting

Harvey's Astra case study describes drafting from court information, firm documents, case-law research, and other matter context. Its memory panel can encode preferences such as prioritizing EDGAR or formatting issues by priority.

The page reports substantial improvements in formatting and context awareness, but gives no task set, comparator models, absolute results, citation-verification rate, or mandatory approval path. More context can produce a better draft. It can also produce a more persuasive error unless the system preserves the relationship between claims, sources, and review decisions.

Together, the cases show useful components. They do not publicly demonstrate one end-to-end evidence record.

A matter-level evidence contract is a versioned record that defines what an AI workflow may use, what it produced, what evidence supports it, who reviewed it, and what allows it to leave the matter file.

It should contain seven modules:

Module Minimum record Failure it exposes
Matter and mandate matter_id, jurisdiction, purpose, allowed tasks, prohibited actions, client instructions, responsible lawyer The agent worked outside the engagement or client instruction
Inputs and permissions Source manifest, snapshot or hash, sensitivity, owner, ethical wall, approved tool and retention mode Matter leakage, stale inputs, or unauthorized reuse
Claim-to-source ledger claim_id, exact passage, citation, authority level, jurisdiction, date checked, support or conflict status A real citation does not support the proposition, or adverse authority is missing
Run provenance Model, search index, plugins, workflow and prompt versions, run time, controlled artifact identifier The result cannot be reproduced after a model or index change
Review decision Reviewer, review scope, method, exceptions, revisions, approval or rejection, timestamp Human review is asserted but cannot be demonstrated
Release and delivery Approved document hash, recipient, channel, time, required AI disclosure A draft is mistaken for the filed or client-delivered version
Correction and closure Error impact, containment, notification, corrected version, re-verification, root cause, closure and reopening rule An error is patched without tracing downstream effects

The contract does not require publishing privileged material or model chain-of-thought. It requires a controlled reference to the evidence that authorized reviewers need.

The useful state machine is simple:

ingested -> generated -> verified -> approved -> delivered
                 |            |
                 +-> rejected +-> reopened after a material change

Each transition needs evidence. A source link moves a claim into the verification queue. It does not move the claim directly to approved.

Review should scale with risk

ABA Formal Opinion 512 is the strongest general professional baseline here. It says lawyers should understand the capabilities and limitations of the tools they use, protect client information, supervise internal and external assistants, review analysis and citations before tribunal submissions, and remain responsible for client work.

The opinion does not impose one review method for every task. Review intensity depends on the tool, task, prior validation, and risk. That supports a risk-based gate:

Work type Practical gate
Formatting, classification, internal routing Automated checks plus documented sampling
Research memo or contract draft Claim-level source checks, jurisdiction and authority review, named reviewer
Court filing, client advice, securities disclosure, external commitment Complete independent verification of material claims, explicit responsible-lawyer approval, frozen delivery version

The distinction matters. A mandatory full review of every low-risk transformation destroys the efficiency gain. A sampling policy applied to a court filing turns efficiency into unmanaged risk.

NIST's Generative AI Profile supports the operational side: system inventories, version history, provenance, human-oversight roles, third-party controls, incident response, and rollback. These are voluntary cross-industry recommendations, not new legal duties. They help translate professional responsibility into system fields.

Measure the complete review system

Once the evidence contract exists, a firm can measure the workflow that actually produces legal work:

  • Claim traceability: proportion of material claims linked to an exact source passage.
  • Authority verification: proportion checked for jurisdiction, status, date, and proposition support.
  • Adverse-authority coverage: required conflicts found, resolved, or explicitly left open.
  • Review yield: claims changed, rejected, or escalated by human review.
  • Change propagation: affected sections correctly reopened after a fact, term, or authority changes.
  • Delivery integrity: delivered artifact matches the approved version and required disclosure.
  • Correction latency: time from error discovery to containment, notification, re-verification, and closure.

These metrics reveal a useful paradox. A review process that changes nothing may be efficient, or it may be ceremonial. Review yield, exception quality, and downstream corrections help distinguish the two.

This is also where the article differs from a general enterprise AI governance framework. Organization-level policies define who owns AI risk. The matter-level contract records whether one specific legal deliverable earned release.

Frequently asked questions

What is Astra for Law?

It is OpenAI's legal configuration of GPT-6 Astra with a legal search index, specialized instructions, tools, and controls. As of September 25, 2026, it is available to selected U.S. law firms through early access, with API access announced as coming soon.

Does a 54% benchmark score mean Astra for Law is reliable?

It means the complete system passed OpenAI's stated overall correctness check on 54.0% of 200 private validation questions under the tested setup. It does not measure production error rates, client outcomes, contract drafting, filings, or unsupervised use.

Why cannot that score be compared with the Vals leaderboard?

OpenAI used the 200-question private validation split. Vals' public leaderboard uses a separate 208-question test split. Different splits and potentially different configurations make direct ranking invalid.

At minimum: matter authorization, source and permission records, claim-to-passage links, authority status, model and tool versions, reviewer decisions, approved delivery version, corrections, and closure conditions.

How should lawyers verify AI-generated citations?

Verification should confirm that the source exists, remains valid, belongs to the correct jurisdiction, supports the exact proposition, fits the client's facts, and does not omit controlling adverse authority. The appropriate scope depends on the task and applicable professional rules.

No. It can improve discovery and passage retrieval. It cannot decide professional responsibility, client objectives, factual fit, strategic judgment, or whether a document is ready to file or deliver.

The next action

Choose one live matter workflow and add the contract before expanding automation. Start with three gates: every material claim has an exact source passage, every external deliverable has a named approver, and every material change reopens affected claims.

That turns legal AI from a fluent drafting component into a system whose work can be inspected, challenged, corrected, and responsibly delivered.

References


Comment