从 Modal 的 Kimi K2.6 案例提炼一套负载优先的推理调优流程:先测请求形态,再依次处理解码、缓存和路由瓶颈。
A workload-first playbook for Kimi K2.6 inference: measure request shape, find the current bottleneck, then tune decode, cache, and routing.
DeepSeek V4.1 用 SWA 有界重放压缩持久 KV 缓存。本文拆解缓存命中时变化的状态、已有证据与上线前测试矩阵。
DeepSeek V4.1 cuts persistent KV storage with bounded replay. See what changes on a cache hit, what was tested, and how to validate it.
KV cache 操作让预训练模型并发观察、思考和响应成为可能。机制已有实验,生产运行时仍在形成,状态治理将成为新的系统边界。
KV-cache manipulation lets pretrained LLMs observe, reason, and respond concurrently. The mechanism works; the production runtime is still emerging.
一套可复用的 Qwen3.8-27B 实战评测合同:比较 tok/s 之前,先控制量化、引擎、缓存、上下文、推理强度、质量与并发。
A practical Qwen3.8-27B benchmark contract that controls quantization, engine, cache, context, reasoning effort, quality, and concurrency before compa