Administrator
Published on 2026-10-04 / 9 Visits
0
0

Kimi K2.6 Inference Tuning Starts With Workload Shape

Kimi K2.6 inference tuning starts with the shape of production traffic, not a list of model optimizations. Modal's coding-agent workload combined roughly 100,000 input tokens, 500 output tokens, repeated prefixes, and multi-turn sessions. That shape moved the bottleneck from decode bandwidth to KV-cache capacity and then to state-aware routing.

Reading time: 9 minutes · Length: about 1,900 words

TL;DR

  • Modal reports 2.8x per-user interactivity and 5.6x aggregate per-replica throughput for one Kimi K2.6 serving stack.
  • Those figures belong to B200 GPUs, SGLang, and a coding-agent workload with an input-to-output ratio near 200:1.
  • The useful pattern is a sequence: profile the workload, fix decode latency, recover KV capacity, then make routing cache- and load-aware.
  • Every successful optimization moves the bottleneck, so the next step must come from a new measurement.
  • Speculative decoding can preserve the target distribution; cache or weight quantization requires application-quality evaluation.

Start with five workload measurements

An inference benchmark without a workload contract is hard to reuse. Tokens per second can change when output length changes. Speculative acceptance changes between code and prose. Cache hit rate changes when users edit earlier messages or when a router moves sessions.

Before changing an engine, collect five distributions from production or a privacy-safe replay:

Signal Why it changes the design
Input tokens per request Drives prefill cost and KV footprint
Output tokens per request Drives sequential decode time
Prefix overlap across turns Determines the value of prefix/KV reuse
Concurrent requests per session Determines whether strict affinity creates hot replicas
Active sessions per replica Connects latency, throughput, and cache pressure

Modal's Kimi K2.6 engineering retrospective describes a core workload around 100k input tokens and 500 output tokens, a ratio of 200:1. A session repeatedly sends its accumulated history, so later turns share most of their prefix with earlier turns.

That traffic is structurally different from short chat, batch summarization, or long-form generation. It rewards fast decode for interactivity, but its aggregate cost is dominated by long inputs and the state needed to avoid recomputing them.

Freeze a baseline before choosing a technique

Use at least four service-level metrics:

  • TTFT: time to first token, including queueing and prefill;
  • inter-token latency or output TPS: user-visible decode speed;
  • tokens per minute per GPU: aggregate cost-performance;
  • p95 end-to-end latency: the tail users actually encounter.

Add cache hit rate, KV utilization, queue depth, and load variance per replica as diagnostic metrics. Segment every chart by input length, output length, concurrency, and prefix-reuse band.

The baseline should use frozen traces for repeatability and a recent production sample for realism. Modal notes that nearly every metric in its project depended on both system configuration and workload data. Controlled traces explain a change; production-shaped traces test whether the explanation survives reality.

Stage 1: remove the decode-bandwidth bottleneck

Modal first optimized interactivity. During autoregressive decode, each step must access large model state to produce the next token. On its deployment, HBM bandwidth was the initial latency bottleneck.

The team combined tensor parallelism with a custom DFlash speculative decoder. Speculative decoding uses a cheaper drafter to propose several tokens, then validates them in parallel with the target model. The validation step preserves the target model's output distribution when implemented correctly. NVIDIA's technical introduction explains the rejection process: accepted draft tokens remain, and generation falls back to the target model at the first rejection.

The optimization works when each expensive target-model step accepts enough proposed tokens to amortize drafting and verification. Its leading indicators are:

  • accepted tokens per step;
  • drafter time as a fraction of target time;
  • output TPS at fixed concurrency;
  • quality parity against the same target configuration.

Modal fine-tuned its drafter on representative coding traces. It reports average acceptance length increasing from 5.00 to 5.84 tokens per step, producing an incremental 20% speedup. The detail matters: a generic drafter may leave performance on the table because acceptance depends on the workload.

The stopping condition for this stage is not maximum output TPS in isolation. Stop when per-user interactivity reaches the product target and further parallelism reduces per-GPU efficiency or creates a more expensive bottleneck.

Stage 2: give the KV cache the recovered budget

Faster decode exposed the next constraint: concurrency was limited by HBM capacity.

The official Kimi K2.6 model card describes a one-trillion-parameter MoE model with 32 billion activated parameters, 61 layers, a 256K context window, MLA attention, and native INT4 quantization. Modal served it on B200 GPUs in NVFP4 and estimated roughly 595 GB of weights in its deployment.

Modal compared four- and eight-GPU tensor-parallel replicas. Its capacity arithmetic estimated about 0.45 million cached tokens for TP4 and 3.0 million for TP8. The larger replica offered six times the KV headroom, but the measured per-GPU performance did not justify it under the interactivity target. The team kept TP4 and attacked memory use directly.

That decision illustrates the workload-first rule. More memory and more GPUs are capabilities. They become improvements only when the constrained objective moves.

The capacity work had three layers:

  1. Remove redundant intermediate allocations and return HBM to the cache.
  2. Reduce precision where task-level evaluations permit it.
  3. Add a slower CPU-memory cache tier to soften overload rather than falling directly into full recomputation.

Modal reports that quantizing the KV cache from BF16 to FP8 doubled token capacity. It also quantized shared experts and reduced peak draft-model intermediate memory. In its evaluations, these changes did not create meaningful quality degradation beyond run-to-run nondeterminism.

Treat that statement as environment-specific evidence. Unlike distribution-preserving speculative decoding, quantization can change application outcomes. Re-run coding, tool-use, long-context, and safety evaluations on your exact serving build.

The operational metric for this stage is not nominal context length. It is how cache hit rate behaves as concurrency and session length rise. A large advertised context window does not guarantee that a replica can retain several active histories without eviction.

Stage 3: route requests to state, not only servers

After the single replica improved, the bottleneck moved to multi-replica request placement.

A KV cache makes replicas stateful for performance, even though a miss does not break correctness. Session affinity is the obvious first step: hash a session ID so later turns return to the replica holding its prefix. It improves locality but leaves three failure modes.

First, one session may issue concurrent requests and overload its assigned replica. Second, sessions have unequal work; counting session IDs ignores differences between a 3k-token turn and a 300k-token turn. Third, adding replicas can remap warm sessions and trigger cold prefills.

Modal's revised routing layer addressed each failure:

  • split highly concurrent sessions after a load threshold;
  • place new sessions using running requests and KV utilization, then retain affinity;
  • prefer new sessions on new replicas during scale-out to reduce cache relocation.

The team reports replica-load variance below half the mean in a sample deployment, steadier tail TTFT, and per-replica throughput closer to the optimized single-replica result.

The general rule is that locality and balance are competing objectives. Perfect locality can create a hot node. Perfect balance can destroy warm state. A useful router optimizes the cost of the next placement, using both load and reusable state.

A bottleneck migration runbook

The case becomes transferable when rewritten as a loop rather than a recipe.

1. Define the constrained objective

Choose a product target such as p95 TTFT below a threshold while maintaining minimum per-GPU throughput. Without the constraint, a benchmark can improve one axis by silently sacrificing another.

2. Reproduce the problem on frozen traces

Capture representative length, overlap, and concurrency bands. Store engine version, model revision, precision, parallelism, hardware, scheduler settings, and router policy with every result.

3. Identify the active resource limit

  • High decode latency with spare compute suggests memory-bandwidth or sequential-step limits.
  • Falling cache hit rate and recomputation at higher concurrency suggest KV-capacity pressure.
  • Good single-replica results with bad cluster tails suggest placement or autoscaling problems.

4. Change one layer

Apply a targeted intervention: speculative decoding, allocation cleanup, cache precision, an extra cache tier, or routing policy. Keep the workload and other controls fixed.

5. Re-run quality and performance gates

Compare both central tendency and tails. Run target-task quality evaluations for any lossy change. A throughput gain that changes tool calls or patch correctness fails the service contract.

6. Locate the new bottleneck

Once the primary constraint moves, stop optimizing the old layer. The next measurement chooses the next intervention.

Common mistakes

Copying the 2.8x and 5.6x figures into a capacity plan. They are author-reported results for a specific stack and traffic shape, not portable constants.

Using average request length. Averages hide the tails that exhaust KV capacity and create p95 latency.

Maximizing cache locality. Affinity needs a load escape hatch for concurrent or unusually large sessions.

Treating quantization as a free optimization. Some changes affect only performance behavior; others can alter model behavior. The latter require application evaluations.

Comparing engines with different workloads. Changing output length, prefix overlap, or concurrency can reorder configurations without any engine change.

FAQ

Should I optimize TTFT or output TPS first?

Start from the product constraint. Long-prefill workloads may be TTFT-bound; interactive coding agents also need adequate decode TPS. Track both, then focus on the one violating the user-facing target.

When does speculative decoding help most?

It helps when draft tokens have a high acceptance rate and the drafter is much cheaper than the target. Acceptance is workload-dependent, so code traces and prose can produce different results.

Why does KV-cache performance collapse with concurrency?

More active histories compete for finite HBM. Evictions force expensive prefix recomputation, reducing both interactivity and throughput.

Is session affinity enough for agent workloads?

It is a useful baseline, but concurrent requests, uneven session sizes, and scale-out remapping can create hot replicas or cold prefills. Add load and KV state to placement decisions.

Will Modal's configuration work for another model?

The diagnostic order may transfer. The winning parallelism, precision, cache tiers, and router thresholds must be re-measured for the new model, hardware, engine, and traffic.

References

Next action: export one hour of request metadata, bucket it by input length, output length, prefix overlap, and session concurrency, then rerun the current serving baseline. That profile determines whether decode, cache, or routing deserves the next engineering cycle.


Comment