Administrator
Published on 2026-09-23 / 13 Visits
0
0

Audit GPT-6 Prompt Caching: 9 Ways to Miss

GPT-6 prompt caching is a controllable API contract. Applications can choose implicit or explicit breakpoints, prewarm reusable prefixes, compare requests with cache diagnostics, and measure cache writes separately from reads. The engineering task is to prove that a real request family reuses the intended prefix and saves total cost without becoming dependent on warm cache state.

Reading time: 9 minutes · About 1,850 words

TL;DR

  • GPT-5.6 and later, including GPT-6, require at least 1,024 visible input tokens in a reusable prefix.
  • Cache writes cost 1.25 times the uncached input rate; reads cost 0.1 times that rate.
  • Explicit-only mode performs no cache read or write when the request contains no explicit breakpoint.
  • Stable content belongs before a breakpoint. User-specific, time-sensitive, and frequently changing content belongs after it.
  • Keep model, service tier, tools, schemas, formats, request-level reasoning effort, verbosity, and earlier input stable.
  • Diagnostics report only the first classified miss reason. Fix it, repeat the comparison, and continue until the usage evidence matches the design.
  • Warm cache is a performance optimization. Every request must remain correct on a cold miss.

First separate three different cache problems

OpenAI's prompt caching guide describes cross-request reuse of a matching prompt prefix. The service preserves key-value tensors for an eligible prefix and later requests can reuse that computation.

This API feature has a different boundary from two related topics already covered here:

GPT-6 prompt caching exposes a request contract: which rendered prefix is eligible, where lookup can occur, what changed, and how usage is billed. It does not provide application memory and should not become a correctness dependency.

The cost model changed the break-even point

For GPT-5.6 and later, OpenAI documents a cache-write price of 1.25 times the ordinary uncached input-token rate and a cache-read price of 0.1 times that rate. One full write followed by one full reuse therefore costs 1.35 units of ordinary input processing, compared with 2 units for two cold requests. One write plus nine reads costs 2.15 units, compared with 10 cold units.

For a reusable prefix with ordinary input cost C, expected prefix cost is approximately:

cached plan = 1.25C × writes + 0.10C × reads + miss cost
cold plan   = 1.00C × requests

This makes reuse frequency part of the design. A prefix written once and never reused costs more than processing it cold. A stable policy shared across many requests can repay the write quickly.

The minimum cacheable prefix is 1,024 visible input tokens. OpenAI-provided hidden system tokens do not count toward that minimum. Cache entries for GPT-5.6 and later remain available for at least 30 minutes after the latest write or reuse under the current 30m setting. Treat that as a minimum availability window rather than a hard expiration timestamp.

Choose implicit or explicit-only mode deliberately

In implicit mode, OpenAI places a breakpoint at the latest eligible message. It is a practical default for conversations that preserve history and append new turns.

In explicit-only mode, the application marks the reusable boundary:

response = client.responses.create(
    model="gpt-6-sol",
    prompt_cache_options={"mode": "explicit"},
    input=[
        {
            "role": "developer",
            "content": [{
                "type": "input_text",
                "text": stable_policy,
                "prompt_cache_breakpoint": {"mode": "explicit"},
            }],
        },
        {"role": "user", "content": changing_request},
    ],
)

Explicit-only mode creates no cache read or write when the request contains no explicit breakpoint. Content after the last selected breakpoint is processed at the uncached rate without a cache-write charge. That is useful when a long stable policy precedes a rapidly changing user payload.

Each request can create up to four cache writes. Lookup also has boundaries. In explicit-only mode, OpenAI checks the first two and latest 50 explicit breakpoints. Implicit mode additionally checks its implicit breakpoint, up to 20 earlier eligible message endings, and the end of the initial consecutive developer-message block.

The practical rule is simple: use few breakpoints that correspond to real change frequencies. For example, one after a versioned global policy and another after a customer-specific knowledge package. More markers add complexity without guaranteeing useful reuse.

Nine cache-miss reasons form an audit matrix

Prompt cache diagnostics compares a current request with a recent completed response from the same organization. Set prompt_cache_options.comparison_response_id to the baseline response ID, then inspect prompt_cache_diagnostics and the response usage.

Diagnostic reason Contract drift Preferred fix
model_changed A different model served the request Pin the model for the request family
prompt_cache_key_changed Accounting or isolation key changed Keep a stable key within the intended group
service_tier_changed Processing tier changed Keep tier stable for comparable requests
tools_changed Tool names, order, descriptions, or schemas changed Preserve the tools array; control availability separately
text_format_changed Structured-output instructions or schema changed Version formats and group requests by schema
reasoning_effort_changed Request-level effort changed hidden instructions Use GPT-6 configuration_update for mid-conversation changes
verbosity_changed Response-detail instructions changed Keep verbosity stable within the family
context_compacted Earlier history was replaced Establish a new baseline after compaction
input_changed Earlier messages, instructions, IDs, or timestamps changed Move dynamic content after the breakpoint and append turns

Diagnostics classify the first reason they can identify. A request can contain several drifts. Fix the reported cause, run the same comparison again, and repeat. unavailable is not evidence of a hit, and an expired diagnostic record can return comparison_response_not_found.

Usage fields remain the billing and reuse evidence. Check usage.input_tokens_details.cached_tokens and cache-write tokens rather than treating the diagnostic label as a cost report.

Keep tools and reasoning append-only

Tool definitions sit inside the rendered prefix. Renaming one function, reordering tools, editing a description, or changing a JSON Schema can invalidate reuse.

When the application needs different tool availability, preserve the full tools list and use tool_choice: "none" or allowed_tools rather than removing definitions. Deferred tool search and additional_tools can append newly discovered tools after the existing prefix.

GPT-6 also supports a configuration_update input item for changing reasoning effort during a conversation:

{
  "type": "configuration_update",
  "reasoning": { "effort": "high" }
}

Keep the request-level reasoning.effort at its original value. Rewriting that top-level setting can alter earlier hidden instructions and reduce reuse.

Prewarm only when expected reuse pays for it

prompt_cache_options.prewarm: true prepares an eligible prefix without generating output. It can lower time to first token for a known upcoming workload, such as loading a shared policy before a scheduled evaluation batch.

Prewarm writes are billed at the normal cache-write rate. The same break-even logic applies. Prewarming an uncertain or one-off request converts possible latency into certain cost.

Define a prewarm trigger with four fields:

  1. expected request count;
  2. maximum time until first use;
  3. prefix version;
  4. cancellation behavior when the workload disappears.

A repeatable deployment audit

Before enabling caching for a request family, freeze ten to fifty representative requests and run this sequence:

  1. Cold baseline: record input tokens, output tokens, p50 and p95 time to first token, total latency, and task success.
  2. Write: send the stable prefix with the intended implicit or explicit breakpoint and record cache-write tokens.
  3. Reuse: replay requests with only the designed suffix changing. Record cached-token ratio and cost.
  4. Mutation tests: change model, tier, tool order, schema, verbosity, effort, one early instruction, and compaction state one at a time.
  5. Diagnostics: compare each mutation with the baseline and save the reason.
  6. Cold-correctness test: run the same tasks after a deliberate miss and verify identical task-level acceptance criteria.
  7. Production gate: alert on declining cached-token ratio, unexpected writes, latency regression, and higher cost per completed task.

The Prompt Caching Dashboard in the OpenAI Platform helps monitor aggregate behavior. Keep request-level traces for root-cause analysis because a fleet-wide hit-rate average can hide one customer or prompt version that constantly rewrites its prefix.

FAQ

Is GPT-6 prompt caching automatic?

Implicit mode places an eligible breakpoint automatically. Explicit-only mode requires at least one explicit breakpoint; without one, the request performs no prompt-cache read or write.

Does a breakpoint need 1,024 tokens before it?

The reusable visible prefix must meet the 1,024-token minimum for GPT-5.6 and later. Hidden OpenAI system content does not count toward that threshold.

Does GPT-6 require prompt_cache_key?

No. On GPT-5.6 and later it is optional and is mainly useful for separate cache accounting or isolation by customer, user, or workspace. OpenAI handles cache routing automatically on these models.

Why can tool reordering cause a miss?

OpenAI caches the rendered prefix, including tool definitions and order. A structural change before the breakpoint changes that prefix.

Can an application clear the prompt cache manually?

OpenAI's current guide does not expose a manual purge operation. Version the prefix or accounting key when logical isolation is required, and design every request to work on a cold cache.

Are diagnostics billed separately?

The diagnostic feature has no additional fee and no separate rate-limit charge. Any extra baseline or retry Responses API calls are billed normally.

Make the cache observable, then optional

Take one high-volume request family and write down its stable prefix, dynamic suffix, breakpoint, expected reuse count, isolation key, and cold-cache acceptance test. Run the seven-step audit before rollout. A useful cache reduces measured latency and total cost. A reliable application remains correct when that cache is absent.

References


Comment