An audit of OpenAI Jalapeño's first benchmarks: what the 1.5–1.9x performance-per-watt lead proves, what it omits, and why full-stack inference is the