Administrator
Published on 2026-09-13 / 26 Visits
0
0

DeepSeek V4.1 SWA Bounded Replay: What Cache-Hit Reconstruction Forgets

DeepSeek V4.1 SWA Bounded Replay makes a precise trade: keep less local attention state, then reconstruct an approximation from the last 128 prompt tokens when that state is missing. DeepSeek reports a global KV footprint of 890 bytes per token and persistent KV storage around one eighth of V4-Flash. Those savings are real. The word approximate decides whether they become a production win.

Reading time: 11 minutes · Word count: approximately 2,200

TL;DR

  • The 890 bytes per token figure describes global KV kept in HBM. It is separate from short-lived sliding-window KV and long-lived persistent cache.
  • Exact sliding-window reconstruction would replay a dependency chain that grows across layers. Bounded replay processes only the latest 128 tokens.
  • The model still has global sparse attention over long context. Bounded replay does not reduce the advertised context window to 128 tokens.
  • What disappears is part of the transitive local-attention contribution from before the replay boundary. DeepSeek explicitly says the reconstructed state is not mathematically identical to full prefill.
  • SGLang reports strong prefill-throughput gains and one paired AIME result with unchanged pass@1. The public evidence remains narrow, so deployment requires boundary-focused tests.

Three cache numbers describe three different constraints

The launch summary is easy to compress into one headline, but the DeepSeek technical report describes separate storage domains.

State Location and lifetime Published result Main mechanism
Global KV Always resident in HBM during runtime 890 bytes per token, around one quarter of V4-Flash CSA2 cross-layer reuse plus FP4 main KV
SWA KV Runtime local state per layer; short-lived host pool for reuse Removed from long-lived persistent storage 128-token window plus bounded replay on a host-pool miss
Persistent KV SSD or host memory for prefix reuse Around one eighth of V4-Flash Global KV compression multiplied by removal of persistent SWA KV

The one-eighth figure is therefore not a new per-token HBM number. The report attributes it to two multiplicative changes. Persisted global KV is about one quarter of the previous footprint, while SWA KV, previously close to half of persistent capacity, is no longer retained in the long-lived cache.

DeepSeek says V4 kept both types of persistent state for more than 72 hours under typical workloads. V4.1 keeps global KV for at least 72 hours, but places SWA KV in a distributed pool using 10% of host DRAM per machine and gives it a lifetime measured in minutes. A miss in this short-lived pool activates encoder-side bounded replay.

Simple arithmetic makes the scale visible. At 890 bytes per token, a one-million-token sequence carries about 890 MB, or 849 MiB, of global KV in aggregate. Eight such live contexts imply about 7.12 GB before allocator overhead and other state. Per-GPU usage depends on tensor parallelism, sharding, batching, and the serving engine. The arithmetic validates the memory order of magnitude, not a single-card deployment promise.

The bottleneck has moved. HBM and persistent-storage pressure fall, while reconstruction compute, cache scheduling, and tail latency gain importance.

Why exact replay grows across layers

DeepSeek V4.1-Flash has 40 Transformer layers arranged as a 20-layer causal encoder and a 20-layer decoder. Each layer has a sliding-window attention branch with a window of 128 tokens. Most layers also have a compressed global-attention branch through CSA2.

A 128-token local window at one layer does not imply a 128-token effective dependency after many layers. A token can receive information from an earlier neighbor, then pass that mixed representation to another token at the next layer. Local dependencies accumulate through the stack. DeepSeek states that exact reconstruction of sliding-window KV across L layers would require replaying approximately L × n_win tokens.

Bounded replay caps that work at n_win, which is 128 for V4.1-Flash. If replay starts at position s, a query at position i only sees local keys from max(s, i-W+1) through i. The replay boundary becomes an artificial beginning for the local-attention path.

DeepSeek applies this idea in two places.

Encoder replay handles a persistent-cache hit without SWA state

The prefix cache retains compressed global KV and indexer keys. When global KV hits but short-lived encoder SWA KV is missing, the engine replays the cached prefix's final 128 tokens together with the uncached suffix. Replayed prefix tokens regenerate only SWA KV. Cached global KV is reused without recomputation or overwrite, while suffix tokens produce both global and local state.

This makes persistent prefix caching depend only on global KV. It also means that state for the new suffix depends on where the cache hit occurred. DeepSeek explicitly notes that suffix global KV and SWA KV are not mathematically identical across different hit positions.

Decoder replay avoids a full upper-stack prefill

Under the causal encoder-decoder design, decoder global KV is projected from the final encoder hidden states. Decoder layers still need their own local SWA state for the first decode steps. Instead of running the complete prompt through every decoder layer, bounded replay sends only the prompt's final 128 encoder outputs through the decoder stack under the same truncated window.

The reconstructed decoder SWA KV is used for decoding and excluded from prefix caching. DeepSeek also simulated the replay path during post-training so that the model could adapt to the approximation.

What does reconstruction actually forget?

The 128-token replay boundary does not erase the entire earlier prompt. CSA2 still provides compressed global attention over distant positions, selecting up to 512 entries for a query. V4.1 still advertises a context window of up to one million tokens.

The missing information is more specific: the complete transitive contribution that earlier tokens would have carried into the current local window through layer-by-layer SWA propagation.

Consider the first token inside a replay segment. During full prefill, its hidden state could attend to local predecessors outside the segment. Those predecessors already contain mixtures from their own local neighborhoods at earlier layers. During bounded replay, the first token has no local keys before s. That missing local contribution then affects later replayed tokens. Global sparse attention can recover relevant distant content through a different path, but it does not make the reconstructed local state mathematically identical.

This creates three distinct failure surfaces:

  1. Boundary sensitivity: the same suffix can receive different approximate state when the cache-hit position changes.
  2. Local-chain tasks: tasks that depend on fine-grained transformations propagated across a boundary may be more sensitive than tasks served well by global retrieval.
  3. Error accumulation: a small representation difference can influence long generation, repeated tool calls, or later cache entries even when the first answer looks correct.

The PowerAttention paper, cited by DeepSeek, argues that practical effective receptive fields can be smaller than the theoretical stacked-window range. That provides a design motivation for bounded replay. It does not establish equivalence for every V4.1 workload.

What the public evidence proves

DeepSeek says its experiments found negligible response-quality impact. The technical report does not publish a dedicated bounded-replay ablation table with task mix, cache-hit positions, variance, or confidence intervals. Its limitations section is unusually useful: potential CSA2 selection errors and approximate SWA reconstruction may still degrade capability in untested boundary cases, especially around sparse long-context retrieval and cache resumption.

The SGLang and Miles integration report adds a concrete implementation result. With batches of eight 8K-token prompts, decoder-side replay increased prefill throughput by 1.56 times on eight H200 GPUs and 1.37 times on four GB300 GPUs. On the four-GB300 setup, a paired AIME 2026 evaluation produced 453 correct answers out of 480 with replay disabled and the same 453 out of 480 with replay enabled.

That is meaningful evidence for one mechanism under one configuration. It supports a narrow claim: decoder replay improved prefill throughput without changing aggregate AIME pass@1 in that paired run. It does not prove identical logits, unchanged long-context retrieval, stable Agent behavior, or equal p99 latency.

The implementation boundaries matter too. SGLang exposes encoder and decoder replay as opt-in flags. Encoder replay excludes speculative decoding. Decoder tail-only computation does not support input log probabilities or full prompt hidden-state capture. A team using any of those features already has a concrete compatibility gate.

Release maturity is another gate. As of September 13, 2026, SGLang's DeepSeek V4.1 support pull request remains open against main; replay fixes have landed on its dsv4.1 integration branch. The branch's encoder replay helper runs only for a normal, non-speculative extend, requires an even cached-prefix boundary, and explicitly clears multimodal replay input. A generic install command on the model card is not evidence that a released serving build reproduces DeepSeek's production cache policy. Pin the tested branch or commit and record it with the result.

A production validation matrix

The right comparison freezes the model, serving build, prompt corpus, sampler, and hardware, then changes only bounded replay. Test the workload you intend to run rather than a generic average. Each prompt should traverse four explicit states where the engine supports them: cold full prefill, warm global-and-local hit, global hit with local miss, and the same global hit cut at a different prefix position.

Dimension Minimum variants Why it matters
Cache-hit position Beginning, just before and after a 128-token boundary, late prefix Exposes position-dependent reconstruction
Context length 8K, 128K, 512K, target maximum Separates published 8K evidence from long-context behavior
Uncached suffix Empty, short user turn, long tool output, image plus text Changes how much new state depends on reconstruction
Task Boundary passkey, multi-hop retrieval, code edit, structured extraction, tool loop Probes both local propagation and global retrieval
Concurrency 1, normal load, saturation Reveals replay scheduling and bandwidth contention
Output length Short answer, long generation, multi-turn continuation Detects downstream accumulation

Measure both resource and outcome metrics: global and SWA KV bytes, host-pool miss rate, replayed tokens, prefill throughput, time to first token, end-to-end latency at p50, p95 and p99, tokens per second, task success, structured-output validity, tool-call correctness, and safety failures.

Use paired prompts around deliberately shifted cache boundaries. A stable average can hide a narrow boundary regression. For nondeterministic sampling, compare enough repeated runs to estimate task-level differences. The production decision should be tied to cost per successful task, which is the same discipline described in model routing by task shape.

Write the pass criteria before running the benchmark. A practical release contract can require zero new critical tool or safety failures, a predeclared maximum task-success regression with a confidence interval, p99 latency inside the existing service objective, and positive throughput or capacity gain at representative load. The acceptable regression budget depends on the application. Freezing it in advance prevents a favorable memory chart from moving the quality threshold after the fact.

Start with replay disabled as the reference. Enable decoder replay first because SGLang provides public throughput and quality evidence for that path. Add encoder replay after measuring short-lived SWA hit rates and confirming that speculative decoding is outside the required configuration. Roll back when boundary-task accuracy, tool correctness, or tail latency crosses a predefined threshold.

Frequently asked questions

What is SWA Bounded Replay in DeepSeek V4.1?

It is a deployment optimization that reconstructs missing sliding-window KV by replaying only the prompt's latest 128 tokens. Exact reconstruction would follow a much longer dependency chain across layers.

Does DeepSeek V4.1 forget everything older than 128 tokens?

No. The model retains a compressed global-attention path over long context. The approximation truncates the local SWA reconstruction path at the replay boundary.

Is bounded replay equivalent to full prefill?

No. DeepSeek states that encoder suffix state can depend on the cache-hit position and that reconstructed decoder SWA KV is not mathematically identical to a full decoder forward pass.

How much KV cache memory does DeepSeek V4.1 use?

DeepSeek reports 890 bytes per token for global KV in HBM, around one quarter of V4-Flash. It separately reports persistent KV storage around one eighth of V4-Flash after removing long-lived SWA KV and compressing global KV.

When should a serving team keep replay disabled?

Keep the reference path when required features are incompatible, or when paired tests show material regression in boundary-sensitive tasks, tool behavior, safety, or tail latency. SGLang currently lists speculative-decoding and observability limitations for specific replay modes.

How should bounded replay be tested before production?

Run replay off and on against frozen prompts while varying cache-hit position, context length, suffix type, concurrency, and output horizon. Judge memory and throughput together with task correctness and tail latency.

The correct conclusion is a moved bottleneck

DeepSeek V4.1 demonstrates that long-context serving can retain much less KV state. Its architecture and SGLang's early implementation results make the memory and prefill case credible. Bounded replay achieves part of that result by accepting approximate local state.

The practical decision belongs at the system level. A service limited by persistent-cache capacity can gain substantially. A workload dominated by cold prefixes, unsupported features, boundary-sensitive transformations, or strict tail-latency targets may expose the next constraint. Treat 890 bytes and one eighth as the start of the measurement plan, then verify where the bottleneck moved.

References


Comment