A useful Gemini 3.8 Live benchmark needs two clocks. The first measures whether a voice agent keeps the conversation responsive. The second measures when its reasoning and tools produce a verified result. Combining both into one average hides the exact trade that distinguishes Gemini 3.8 Live from Gemini 3.8 Live Extended Thinking.
Reading time: 9 minutes · About 1,750 words
TL;DR
- Measure time to first meaningful audio and time to verified task completion separately.
- For Extended Thinking,
turnComplete: trueends an utterance, not necessarily the full interaction. The client must followinteraction_statusuntilIDLE. - Benchmark direct dialogue, slow tools, multi-step reasoning, interruptions, network impairment, and session recovery as different task classes.
- Treat conversational fillers as useful only when they are timely, truthful, and followed by progress.
- Publish distributions, failure rates, task success, and cost per successful task. One latency average or one vendor leaderboard score is insufficient.
Two models create two latency budgets
Google positions gemini-3.8-live for immediate dialogue and direct tasks. It positions gemini-3.8-live-extended-thinking for complex planning, background reasoning, and asynchronous tools. This is an architectural distinction, not just a quality setting.
The standard model uses a familiar lifecycle: a completed turn returns the session to idle. Extended Thinking may speak an acknowledgement, mark that utterance complete, continue reasoning or calling tools, and speak again later. Google therefore adds interaction_status: IN_PROGRESS means work may still be running; IDLE means the interaction has actually finished.
A benchmark that stops its timer at the first audio frame rewards a quick filler even when the answer is late or wrong. A benchmark that records only final completion penalizes a system that keeps the user informed while doing legitimate work. Both measurements are necessary.
Define two top-level clocks:
- Conversation latency: from the end of user speech to the first meaningful model audio.
- Outcome latency: from the end of user speech to a verified answer or completed external action.
The first protects turn-taking. The second protects usefulness. Optimizing either in isolation can move the bottleneck into endpoint detection, tool execution, background reasoning, or synthesis.
Decompose the conversation clock
End-to-end voice latency is a chain of delays. Record timestamps at each boundary:
user speech ends
→ client commits audio
→ server recognizes the turn boundary
→ first model audio arrives
→ first meaningful phrase is heard
→ utterance completes
→ interaction becomes IDLE
→ external result is verified
At minimum, capture these metrics:
| Metric | What it diagnoses |
|---|---|
| End-of-speech to first audio, P50/P95/P99 | Perceived responsiveness |
| End-of-speech to first meaningful content | Filler that masks a slow answer |
| False turn-end and missed turn-end rate | Endpoint-detection quality |
| Barge-in to output stop | Interruption responsiveness |
| Recovery after interruption | Whether state remains coherent |
Time to IDLE |
Full background-reasoning lifecycle |
| Tool request, response, and acknowledgement times | External dependency bottlenecks |
| Verified task-completion time | Actual user outcome |
Use a manually labeled or offline-detected acoustic end of speech as the timer origin. Starting at the client's VAD event hides the endpointing delay that the user actually experiences. Google's Live API exposes automatic and hybrid VAD controls, but it does not publish a universal P50 or P95 end-to-end latency target for these two GA model endpoints. Treat silence thresholds and latency budgets as workload-specific SLOs, not vendor constants.
Use P95 and P99 in addition to medians. Voice interaction is unusually sensitive to tail events because silence, overlapping speech, and repeated acknowledgements are immediately visible to users.
A filler is a protocol event, not proof of progress
Extended Thinking can speak a phrase such as Checking that now while it reasons or waits for tools. This can reduce awkward silence, but it creates a new failure mode: conversational responsiveness without operational progress.
Evaluate each intermediate utterance on three properties:
- Timeliness: did it arrive before the silence became disruptive?
- Truthfulness: does it accurately describe current work rather than imply completion?
- Progress linkage: was it followed by a tool call, a status transition, or a final answer within the allowed budget?
Log audio utterances, turnComplete, interaction_status, tool events, and final verification in one trace. A UI that switches to idle on the first turnComplete: true can misrepresent an active Extended Thinking interaction. Google explicitly requires clients to keep listening until interaction_status: IDLE.
Freeze a task-shape matrix
Do not compare both models with one blended prompt set. The selection decision depends on task shape.
| Task class | Example | Primary clock | Expected candidate |
|---|---|---|---|
| Direct dialogue | answer a factual question | first meaningful audio | Live |
| Fast action | read a sensor or toggle a setting | verified action plus acknowledgement | Live |
| Slow tool | query several external services | progress cadence plus completion | Extended Thinking |
| Multi-step reasoning | diagnose logs and propose a fix | verified task success | Extended Thinking |
| Interruption | user changes the request mid-answer | stop and recovery latency | Both |
| Long session | reconnect and resume state | recovery success and state loss | Both |
For Extended Thinking, test low, medium, and high thinking levels. Freeze prompts, audio files, tools, network conditions, and answer checkers before testing. The standard model does not expose the same configurable thinking level, so report the comparison as two product modes rather than pretending it is a single controlled parameter sweep.
Use repeatable audio and controlled tools
A reproducible harness should include recorded speech and scripted tool servers.
The audio corpus should vary accent, speaking rate, hesitations, background noise, code-switching, numbers, names, and overlapping speech. Include turns that end cleanly and turns with long intra-sentence pauses. Replay the same files through both models and retain timestamps from capture through playback.
Tool mocks should expose fixed latency bands, for example 50 ms, 500 ms, 3 seconds, and 10 seconds, plus timeouts and malformed responses. This separates model orchestration from third-party availability. Then repeat a smaller production-like run against real tools.
Test at least three network profiles: stable broadband, mobile jitter with packet loss, and a reconnect. Google documents that Live API connections last around ten minutes. Audio-only sessions without compression are limited to fifteen minutes, while context compression and session resumption change the long-session path. Recovery belongs in the benchmark rather than in an operational footnote.
Score outcomes, not only speech
Google reports strong benchmark results for Extended Thinking, including 82.6 on Artificial Analysis' Speech-to-Speech Quality Index, 68.6% on τ-Voice, 35.1% on Sierra's τ-Voice banking benchmark, and 97.7% on Big Bench Audio. These are useful launch signals. They do not replace a workload-specific test, and the cited figures arrive through Google's launch material and linked benchmark providers.
For your own tasks, define a deterministic or human-auditable checker:
- Did the requested external action happen exactly once?
- Did the answer use the current tool result?
- Did the model preserve constraints after an interruption?
- Did it disclose uncertainty instead of inventing completion?
- Did the session reach
IDLEwithout abandoned work?
Calculate cost per successful task rather than cost per audio minute alone. A faster model that requires repetition may be more expensive. A slower reasoning model may be justified only when its success gain exceeds its added wait and cost.
A practical release contract
Write thresholds before running the comparison. A production contract can include:
- P95 first meaningful audio stays inside the product's conversational budget.
- P95 barge-in stop and recovery remain below a declared threshold.
- No critical action is marked complete before external verification.
- Tool calls are idempotent across reconnects and interruptions.
- Task success improves enough to justify the Extended Thinking outcome latency.
- Timeout, hallucination, and abandoned
IN_PROGRESSrates stay below fixed limits. - The client treats
interaction_status, not the firstturnComplete, as lifecycle authority for Extended Thinking.
Google's model card lists hallucinations and occasional slowness or timeouts as known limitations. Those limitations should become explicit test cases. The card also lists a 128K-token input window and 64K-token output limit for both audio models; capacity limits still require session-level observation because a nominal context window does not guarantee stable long-call behavior.
FAQ
What is the difference between Gemini 3.8 Live and Extended Thinking?
Gemini 3.8 Live targets immediate dialogue and direct work. Extended Thinking adds configurable background reasoning and requires asynchronous, non-blocking function calls. It can emit several spoken updates before the interaction becomes idle.
Is time to first audio enough for a voice-agent benchmark?
No. It measures perceived responsiveness but can reward empty fillers. Pair it with first meaningful content, time to IDLE, verified task completion, and task success.
Why does turnComplete: true need special handling?
For the standard model it closes the turn. For Extended Thinking it can close one utterance while reasoning or tools continue. Read interaction_status and wait for IDLE before treating the interaction as finished.
How should interruptions be tested?
Interrupt at the start, middle, and end of speech and while a tool is running. Measure output-stop latency, duplicate tool actions, state consistency, and whether the revised request becomes authoritative.
Which model should a production voice agent use?
Route by task shape. Use the standard model when immediate turn-taking dominates. Use Extended Thinking when multi-step correctness justifies a longer outcome clock. Validate both against the same frozen workload.
References
- Google: Introducing Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking
- Google AI for Developers: Thinking in the Live API
- Google AI for Developers: Gemini 3.8 Live Extended Thinking
- Google AI for Developers: Live API overview
- Google AI for Developers: Live API capabilities
- Google AI for Developers: Session management
- Google DeepMind: Gemini 3.8 Audio model card
- Related: Realtime transcription contract for voice agents
- Related: OpenAI low-latency voice AI architecture
The useful question is not which model feels faster in a demo. It is which mode keeps the conversation responsive while finishing the right task, under the failures and interruptions that define production voice systems.