Administrator
Published on 2026-09-21 / 16 Visits
0
0

Patch Half-Life: A Rewrite Signal for AI Coding

Summary: AI coding can move the delivery bottleneck from code generation to CI, testing, and maintenance. Anthropic's test-selection service offers a useful case: three successive patches survived 70 days, 29 days, and less than one day before a three-week redesign stabilized the backlog. This article turns that pattern into a measurable rewrite decision without treating one company's experience as a universal threshold.

Reading time: 9 minutes · About 2,100 words

TL;DR

  • Patch half-life is the time from a workaround's release to the recurrence of the same capacity or reliability breach.
  • A shorter life across successive patches is stronger evidence than code age, line count, or architectural taste.
  • Rewrite only when the recurring constraint is understood, the projected repair bill exceeds migration cost, and equivalence plus rollback can be verified.
  • Prefer an incremental replacement over a big-bang cutover when the service is critical.
  • AI lowers implementation cost. It does not remove the cost of discovering behavior, validating outputs, or operating two systems during migration.

Anthropic's CI case shows a moving bottleneck

Anthropic reported that its engineers were shipping eight times as much code per quarter as they had from 2021 through 2025, while its test corpus grew tenfold. CI job volume rose 25 times in six months. Code generation and pull-request review accelerated; test selection became the constraint (Anthropic's CI engineering report).

The service had a stateful listener that collected test results and a selector that used that history to decide which tests should run. Under the new load, a 20-minute listener delay could leave tens of thousands of test updates unapplied. The resulting data was stale, so the selector could schedule irrelevant flaky tests or delay the use of newly added and repaired tests. Anthropic explicitly says this did not mean CI was skipped on pull requests or untested code reached production.

The team applied three fixes:

  1. It doubled the machine's cores. The fix lasted 70 days.
  2. It sharded state by package. The fix lasted 29 days.
  3. It restarted the process daily. That bought less than one day and made the service fall further behind.

The final redesign moved state into an in-memory data store, made listener workers stateless, journaled incoming results, and added a separate rollup consumer. One engineer completed the project in three weeks. Anthropic says the queued-event backlog became flat after cutover and tuning, and the service remained stable at publication time.

These are internal measurements reported by Anthropic, not an independent industry benchmark. They do, however, illustrate a general mechanism also described in Anthropic's broader discussion of Amdahl's law and organizational bottlenecks: speeding up one part of a system moves the limit to the parts that did not speed up.

That mechanism is the point. Faster code production created more work-in-progress for CI. The first patches optimized the overloaded implementation. The redesign removed the stateful singleton that prevented horizontal scaling.

This is the same productivity gap discussed in Zero-Cost Programming: local output can rise while system throughput stays flat. The new question is how to know when another repair has stopped being the rational choice.

Patch half-life is an operational metric, not a metaphorical complaint

Patch half-life is a term proposed in this article. Anthropic does not use it, and it is not a statistical half-life unless a team has enough observations to estimate a decay distribution.

For day-to-day engineering, define patch survival time as:

Sᵢ = time of the next breach of the same service objective − patch i release time

The denominator matters. A patch that survives until an unrelated dependency fails has not expired. A patch expires when the same constrained resource or failure mode crosses the same predeclared boundary again.

Examples of comparable boundaries include:

  • listener lag exceeds 50,000 jobs;
  • p95 queue time breaches the CI service-level objective;
  • memory saturation forces an emergency restart;
  • stale test-selection data exceeds an agreed age;
  • on-call pages from the same failure mode exceed a weekly budget.

Then calculate the survival ratio:

Dᵢ = Sᵢ / Sᵢ₋₁

In Anthropic's reported sequence, the second patch's ratio was about 0.41, and the third was below 0.035. Three observations cannot establish a universal curve. They are enough to show that the same class of intervention was buying sharply less time.

For a larger system, report a rolling median and lower quartile by component and failure class, then normalize against deployment count, traffic, or CI-job volume. Otherwise rising release volume can make patches appear less durable, while improved monitoring can make recurrence appear more frequent. Freeze the failure taxonomy as well: security updates, dependency upgrades, capacity failures, and product-behavior defects should not share one series.

This signal is more useful than statements such as the code is old or the architecture feels messy. Those statements describe discomfort. Patch survival measures whether repair is still changing the system's limiting constraint.

Use five gates before approving a rewrite

A falling patch survival time is a warning, not automatic permission to rewrite. A rewrite should pass five gates.

1. Same-constraint gate

Confirm that successive incidents come from the same binding constraint. Group events by violated objective and causal mechanism, not by alert name.

If one patch fixes memory growth and the next incident is caused by an external API quota, their survival times should not be placed in one series. A declining series built from unrelated failures is measurement theater.

2. Durability gate

Record every material repair as an event:

Field What to record
Boundary The SLO or capacity threshold that triggered work
Patch The smallest intervention intended to restore headroom
Start and expiry Release time and recurrence time for the same breach
Headroom Capacity immediately after the patch and its consumption slope
Human cost Engineering days, on-call interruptions, and investigation time
Risk cost Stale data, missed work, rollback exposure, and customer impact

Two consecutive declines are worth a design review. Three emergency repairs to the same constraint should trigger a formal rewrite comparison. These are review triggers, not universal pass/fail thresholds.

Run that review at the smallest boundary supported by the evidence. A decaying worker or selector may justify a component rewrite; it does not automatically justify replacing the repository, platform, or product.

3. Economic crossover gate

Compare the costs over a fixed decision horizon, such as the next two quarters:

Repair path = expected patches + incidents + delay + accumulated migration debt

Rewrite path = implementation + behavior discovery + dual running + migration + rollback reserve

AI mainly reduces the implementation term. It may also accelerate test creation and migration tooling. It does not erase behavior discovery, production verification, or the cost of running both paths during a safe transition.

Approve the rewrite only when the conservative repair estimate exceeds the conservative replacement estimate. If the result changes when one optimistic assumption is removed, the decision is not ready.

Schedule matters as well. If the conservative survival estimate for the next patch is shorter than rewrite delivery plus a rollback buffer, the patch can only serve as a migration bridge. Treating it as a durable fix creates a predictable period with no protected path.

This follows the structure of Carnegie Mellon SEI's Cost Benefit Analysis Method: compare architectural strategies across cost, benefit, schedule, and risk instead of letting a single technical metric decide.

4. Constraint-removal gate

The target design must remove the measured bottleneck rather than reproduce it with cleaner code.

Anthropic's redesign changed the scaling property: state left the singleton process, listener workers became stateless, and a journal separated ingestion from aggregation. Rewriting the same singleton in a newer framework would have produced a modern-looking version of the same limit.

This is where harness engineering matters. The valuable artifact is the constraint model and its verification interface, not the volume of replacement code.

5. Recoverability gate

No critical rewrite is ready without a way to prove equivalence and retreat safely. At minimum, require:

  • a frozen corpus of representative inputs and expected decisions;
  • contract tests for externally visible behavior;
  • shadow processing or dual reads before authority moves;
  • reconciliation metrics between old and new outputs;
  • a canary that limits blast radius;
  • a rollback route that has been exercised, not merely documented.

Google's canarying guidance adds the operational contract: compare a release candidate with a concurrent control, declare decision metrics in advance, and pause or roll back when they diverge.

The Strangler Fig pattern exists because full cutover rewrites concentrate risk. AWS similarly recommends starting incremental replacement with components that have good test coverage or acute scaling needs (AWS Prescriptive Guidance).

A practical decision table

Observed state Recommended action
Survival time is stable or increasing; failures have different causes Keep patching and improve instrumentation
Survival time shrinks, but the binding constraint is unclear Pause architecture work and improve causal measurement
Same constraint recurs and survival time shrinks twice Fund a time-boxed redesign proof of concept
Repair cost exceeds rewrite cost, but equivalence cannot be tested Build the verification harness before replacing production
Constraint-removing design and rollback path are both proven Migrate incrementally with shadow and canary stages

This table separates three questions that rewrite debates often mix together:

  1. Is repair losing economic value?
  2. Does the proposed design remove the actual limit?
  3. Can the transition be verified and reversed?

A yes to the first question alone is insufficient.

Measure throughput and instability together

A rewrite can make queues look healthy while making releases less safe. The decision therefore needs both throughput and instability metrics.

DORA's software delivery metrics pair change lead time and deployment frequency with change fail rate, failed-deployment recovery time, and deployment rework rate. For a CI or test-selection service, add domain-specific measures:

  • input jobs versus durably recorded results;
  • event backlog age and depth;
  • selector-data staleness;
  • false skips and unnecessary test runs;
  • escaped failures found by periodic full-suite runs;
  • selection-set disagreement during shadow comparison;
  • p95 decision latency;
  • emergency intervention frequency;
  • rollback and reconciliation time.

Google's SRE guidance treats monitoring as the basis for comparing changes over time and making rational operational decisions, while warning that paging logic should remain simple and robust (Monitoring Distributed Systems). That advice is especially relevant when agents use telemetry as their eyes and ears. Bad observability lets an agent optimize a proxy while the real queue grows out of sight.

What this case does not prove

Anthropic's experience does not prove that every team should design for 25-times load, that the third patch always justifies a rewrite, or that AI makes greenfield replacement safe.

It supports a narrower conclusion: when implementation cost falls quickly, the build-versus-repair equation can change before established architecture-review habits catch up. Teams need a time-series signal for repair durability and a verification gate for replacement.

The unit of optimization is the delivery system, not code output. If code becomes cheap while validation remains scarce, generating more code is a non-bottleneck optimization. It creates inventory for the slowest stage.

FAQ

What is the difference between refactoring and rewriting?

Refactoring changes internal structure while preserving behavior in the same evolving system. A rewrite creates a replacement implementation and moves authority to it. In practice, incremental modernization sits between them: replace one boundary at a time while the old system remains operational.

Is three patches a universal rule for rewriting?

No. Three repairs to the same constraint are a useful review trigger because they provide a short time series. The decision still depends on durability, incident cost, architecture, testability, and rollback risk.

Can patch half-life be used for small systems?

Yes, if the team can define a comparable breach. A cron job that repeatedly misses its completion window, a database worker that repeatedly exhausts memory, or a build service whose queue repeatedly crosses its SLO can all produce a survival series.

Does AI make rewrites safer?

AI can make implementation, test generation, and migration tooling cheaper. Safety comes from behavior capture, independent verification, staged authority transfer, and rollback. Faster code changes the economics; it does not change that control structure.

Next action: start the ledger before the next patch

For the next recurring capacity incident, record the violated boundary, release time, restored headroom, expected growth rate, and time to recurrence. After two comparable patches, calculate the survival ratio and price both paths over a fixed horizon.

The rewrite decision should emerge from that ledger. Architecture taste can start the conversation; measured patch durability should decide whether repair still buys time.

References


Comment