KV-cache manipulation lets pretrained LLMs observe, reason, and respond concurrently. The mechanism works; the production runtime is still emerging.