Administrator
Published on 2026-09-16 / 6 Visits
0
0

Gemini 3.8 Live Benchmark: A Dual-Latency Contract for Voice Agents

A useful Gemini 3.8 Live benchmark needs two clocks. The first measures whether a voice agent keeps the conversation responsive. The second measures when its reasoning and tools produce a verified result. Combining both into one average hides the exact trade that distinguishes Gemini 3.8 Live from Gemini 3.8 Live Extended Thinking.

Reading time: 9 minutes · About 1,750 words

TL;DR

  • Measure time to first meaningful audio and time to verified task completion separately.
  • For Extended Thinking, turnComplete: true ends an utterance, not necessarily the full interaction. The client must follow interaction_status until IDLE.
  • Benchmark direct dialogue, slow tools, multi-step reasoning, interruptions, network impairment, and session recovery as different task classes.
  • Treat conversational fillers as useful only when they are timely, truthful, and followed by progress.
  • Publish distributions, failure rates, task success, and cost per successful task. One latency average or one vendor leaderboard score is insufficient.

Two models create two latency budgets

Google positions gemini-3.8-live for immediate dialogue and direct tasks. It positions gemini-3.8-live-extended-thinking for complex planning, background reasoning, and asynchronous tools. This is an architectural distinction, not just a quality setting.

The standard model uses a familiar lifecycle: a completed turn returns the session to idle. Extended Thinking may speak an acknowledgement, mark that utterance complete, continue reasoning or calling tools, and speak again later. Google therefore adds interaction_status: IN_PROGRESS means work may still be running; IDLE means the interaction has actually finished.

A benchmark that stops its timer at the first audio frame rewards a quick filler even when the answer is late or wrong. A benchmark that records only final completion penalizes a system that keeps the user informed while doing legitimate work. Both measurements are necessary.

Define two top-level clocks:

  1. Conversation latency: from the end of user speech to the first meaningful model audio.
  2. Outcome latency: from the end of user speech to a verified answer or completed external action.

The first protects turn-taking. The second protects usefulness. Optimizing either in isolation can move the bottleneck into endpoint detection, tool execution, background reasoning, or synthesis.

Decompose the conversation clock

End-to-end voice latency is a chain of delays. Record timestamps at each boundary:

user speech ends
  → client commits audio
  → server recognizes the turn boundary
  → first model audio arrives
  → first meaningful phrase is heard
  → utterance completes
  → interaction becomes IDLE
  → external result is verified

At minimum, capture these metrics:

Metric What it diagnoses
End-of-speech to first audio, P50/P95/P99 Perceived responsiveness
End-of-speech to first meaningful content Filler that masks a slow answer
False turn-end and missed turn-end rate Endpoint-detection quality
Barge-in to output stop Interruption responsiveness
Recovery after interruption Whether state remains coherent
Time to IDLE Full background-reasoning lifecycle
Tool request, response, and acknowledgement times External dependency bottlenecks
Verified task-completion time Actual user outcome

Use a manually labeled or offline-detected acoustic end of speech as the timer origin. Starting at the client's VAD event hides the endpointing delay that the user actually experiences. Google's Live API exposes automatic and hybrid VAD controls, but it does not publish a universal P50 or P95 end-to-end latency target for these two GA model endpoints. Treat silence thresholds and latency budgets as workload-specific SLOs, not vendor constants.

Use P95 and P99 in addition to medians. Voice interaction is unusually sensitive to tail events because silence, overlapping speech, and repeated acknowledgements are immediately visible to users.

A filler is a protocol event, not proof of progress

Extended Thinking can speak a phrase such as Checking that now while it reasons or waits for tools. This can reduce awkward silence, but it creates a new failure mode: conversational responsiveness without operational progress.

Evaluate each intermediate utterance on three properties:

  • Timeliness: did it arrive before the silence became disruptive?
  • Truthfulness: does it accurately describe current work rather than imply completion?
  • Progress linkage: was it followed by a tool call, a status transition, or a final answer within the allowed budget?

Log audio utterances, turnComplete, interaction_status, tool events, and final verification in one trace. A UI that switches to idle on the first turnComplete: true can misrepresent an active Extended Thinking interaction. Google explicitly requires clients to keep listening until interaction_status: IDLE.

Freeze a task-shape matrix

Do not compare both models with one blended prompt set. The selection decision depends on task shape.

Task class Example Primary clock Expected candidate
Direct dialogue answer a factual question first meaningful audio Live
Fast action read a sensor or toggle a setting verified action plus acknowledgement Live
Slow tool query several external services progress cadence plus completion Extended Thinking
Multi-step reasoning diagnose logs and propose a fix verified task success Extended Thinking
Interruption user changes the request mid-answer stop and recovery latency Both
Long session reconnect and resume state recovery success and state loss Both

For Extended Thinking, test low, medium, and high thinking levels. Freeze prompts, audio files, tools, network conditions, and answer checkers before testing. The standard model does not expose the same configurable thinking level, so report the comparison as two product modes rather than pretending it is a single controlled parameter sweep.

Use repeatable audio and controlled tools

A reproducible harness should include recorded speech and scripted tool servers.

The audio corpus should vary accent, speaking rate, hesitations, background noise, code-switching, numbers, names, and overlapping speech. Include turns that end cleanly and turns with long intra-sentence pauses. Replay the same files through both models and retain timestamps from capture through playback.

Tool mocks should expose fixed latency bands, for example 50 ms, 500 ms, 3 seconds, and 10 seconds, plus timeouts and malformed responses. This separates model orchestration from third-party availability. Then repeat a smaller production-like run against real tools.

Test at least three network profiles: stable broadband, mobile jitter with packet loss, and a reconnect. Google documents that Live API connections last around ten minutes. Audio-only sessions without compression are limited to fifteen minutes, while context compression and session resumption change the long-session path. Recovery belongs in the benchmark rather than in an operational footnote.

Score outcomes, not only speech

Google reports strong benchmark results for Extended Thinking, including 82.6 on Artificial Analysis' Speech-to-Speech Quality Index, 68.6% on τ-Voice, 35.1% on Sierra's τ-Voice banking benchmark, and 97.7% on Big Bench Audio. These are useful launch signals. They do not replace a workload-specific test, and the cited figures arrive through Google's launch material and linked benchmark providers.

For your own tasks, define a deterministic or human-auditable checker:

  • Did the requested external action happen exactly once?
  • Did the answer use the current tool result?
  • Did the model preserve constraints after an interruption?
  • Did it disclose uncertainty instead of inventing completion?
  • Did the session reach IDLE without abandoned work?

Calculate cost per successful task rather than cost per audio minute alone. A faster model that requires repetition may be more expensive. A slower reasoning model may be justified only when its success gain exceeds its added wait and cost.

A practical release contract

Write thresholds before running the comparison. A production contract can include:

  1. P95 first meaningful audio stays inside the product's conversational budget.
  2. P95 barge-in stop and recovery remain below a declared threshold.
  3. No critical action is marked complete before external verification.
  4. Tool calls are idempotent across reconnects and interruptions.
  5. Task success improves enough to justify the Extended Thinking outcome latency.
  6. Timeout, hallucination, and abandoned IN_PROGRESS rates stay below fixed limits.
  7. The client treats interaction_status, not the first turnComplete, as lifecycle authority for Extended Thinking.

Google's model card lists hallucinations and occasional slowness or timeouts as known limitations. Those limitations should become explicit test cases. The card also lists a 128K-token input window and 64K-token output limit for both audio models; capacity limits still require session-level observation because a nominal context window does not guarantee stable long-call behavior.

FAQ

What is the difference between Gemini 3.8 Live and Extended Thinking?

Gemini 3.8 Live targets immediate dialogue and direct work. Extended Thinking adds configurable background reasoning and requires asynchronous, non-blocking function calls. It can emit several spoken updates before the interaction becomes idle.

Is time to first audio enough for a voice-agent benchmark?

No. It measures perceived responsiveness but can reward empty fillers. Pair it with first meaningful content, time to IDLE, verified task completion, and task success.

Why does turnComplete: true need special handling?

For the standard model it closes the turn. For Extended Thinking it can close one utterance while reasoning or tools continue. Read interaction_status and wait for IDLE before treating the interaction as finished.

How should interruptions be tested?

Interrupt at the start, middle, and end of speech and while a tool is running. Measure output-stop latency, duplicate tool actions, state consistency, and whether the revised request becomes authoritative.

Which model should a production voice agent use?

Route by task shape. Use the standard model when immediate turn-taking dominates. Use Extended Thinking when multi-step correctness justifies a longer outcome clock. Validate both against the same frozen workload.

References

The useful question is not which model feels faster in a demo. It is which mode keeps the conversation responsive while finishing the right task, under the failures and interruptions that define production voice systems.


Comment