Migrate realtime examples to RealtimeServiceModeConfig

Pass realtime_service_mode=RealtimeServiceModeConfig() through every
realtime LLM service example (base, async-tool, video, text-output,
persistent-context, update-settings, MCP) so context aggregation uses
the new realtime-mode semantics instead of relying on local VAD as a
workaround.

Where examples previously wired SileroVADAnalyzer into
LLMUserAggregatorParams to coax turn frames out of services that don't
emit them server-side (AWS Nova Sonic, Ultravox, Gemini Live), the local
VAD is now removed. realtime_service_mode keeps context writes correct
without it, and the Phase 1.5 server-side InterruptionFrame fixes for
Nova Sonic and Ultravox keep the bot from talking past the user when
they barge in.

Transcript-logging event handlers move from on_user_turn_stopped /
on_assistant_turn_stopped to on_user_message_added /
on_assistant_message_added, which carry the finalized text in realtime
mode (the turn-stopped events fire before the message is finalized, so
their `content` is None in that mode).

For services that don't emit user-turn frames (Gemini Live, AWS Nova
Sonic, Ultravox) the example now carries a Tier 1 comment block that
spells out which downstream processors won't activate, how to add local
VAD if needed, and the caveat that locally-generated turn boundaries
are a heuristic that may diverge from server-side ground truth.

Adds examples/realtime/realtime-openai-local-vad.py, a new variant of
the OpenAI Realtime example that disables OpenAI's server-side turn
detection and drives turn boundaries locally — useful when you want a
turn analyzer like LocalSmartTurnV3 to decide when the user is done
speaking. Server-emitted turn frames are still preferred when available.

The Gemini Live local-VAD variant already existed; it's been updated in
place rather than rewritten.
This commit is contained in:
Paul Kompfner
2026-05-20 15:51:18 -04:00
parent 20d9bf4af6
commit bff741a647
35 changed files with 537 additions and 158 deletions

View File

@@ -4,6 +4,29 @@
# SPDX-License-Identifier: BSD 2-Clause License
#
"""Gemini Live with locally-driven turn detection.
By default Gemini Live drives the conversation with its own server-side VAD
(see `realtime-gemini-live.py`). That setup doesn't surface
``UserStartedSpeakingFrame`` / ``UserStoppedSpeakingFrame``, so pipeline
processors that depend on those frames (RTVI client speech events,
``TurnTrackingObserver``, ``AudioBufferProcessor`` turn recording,
``UserIdleController``, user mute strategies, voicemail detector) don't
activate.
This variant disables Gemini Live's server-side VAD
(``GeminiVADParams(disabled=True)``) and instead drives turn boundaries
locally with ``SileroVADAnalyzer`` wired into the user aggregator. Use this
variant if you need those downstream processors, or if you want a turn
analyzer like ``LocalSmartTurnV3`` to decide when the user is done speaking.
Caveat: locally-generated turn boundaries are a heuristic and may not match
the provider's actual server-side turn decisions, which is what really
drives the conversation. The two can drift apart in subtle, hard-to-debug
ways, especially around interruptions and overlapping speech. Prefer
server-emitted turn frames (i.e. the base `realtime-gemini-live.py` example)
unless you have a specific reason to drive turn detection locally.
"""
import os
@@ -20,6 +43,7 @@ from pipecat.processors.aggregators.llm_response_universal import (
AssistantTurnStoppedMessage,
LLMContextAggregatorPair,
LLMUserAggregatorParams,
RealtimeServiceModeConfig,
UserTurnStoppedMessage,
)
from pipecat.runner.types import RunnerArguments
@@ -72,6 +96,7 @@ async def run_bot(transport: BaseTransport, runner_args: RunnerArguments):
)
user_aggregator, assistant_aggregator = LLMContextAggregatorPair(
context,
realtime_service_mode=RealtimeServiceModeConfig(),
user_params=LLMUserAggregatorParams(
vad_analyzer=SileroVADAnalyzer(),
),
@@ -107,14 +132,17 @@ async def run_bot(transport: BaseTransport, runner_args: RunnerArguments):
logger.info(f"Client disconnected")
await task.cancel()
@user_aggregator.event_handler("on_user_turn_stopped")
async def on_user_turn_stopped(aggregator, strategy, message: UserTurnStoppedMessage):
# The *_message_added events fire when messages are written to context
# and carry the finalized content. In realtime mode the turn-stopped
# events fire before the message text is finalized.
@user_aggregator.event_handler("on_user_message_added")
async def on_user_message_added(aggregator, message: UserTurnStoppedMessage):
timestamp = f"[{message.timestamp}] " if message.timestamp else ""
line = f"{timestamp}user: {message.content}"
logger.info(f"Transcript: {line}")
@assistant_aggregator.event_handler("on_assistant_turn_stopped")
async def on_assistant_turn_stopped(aggregator, message: AssistantTurnStoppedMessage):
@assistant_aggregator.event_handler("on_assistant_message_added")
async def on_assistant_message_added(aggregator, message: AssistantTurnStoppedMessage):
timestamp = f"[{message.timestamp}] " if message.timestamp else ""
line = f"{timestamp}assistant: {message.content}"
logger.info(f"Transcript: {line}")