从 Modal 的 Kimi K2.6 案例提炼一套负载优先的推理调优流程:先测请求形态,再依次处理解码、缓存和路由瓶颈。
A workload-first playbook for Kimi K2.6 inference: measure request shape, find the current bottleneck, then tune decode, cache, and routing.