A workload-first playbook for Kimi K2.6 inference: measure request shape, find the current bottleneck, then tune decode, cache, and routing.