Bidirectional Streaming Models
A BidiAgent talks to a realtime model: a hosted speech-to-speech service, such as
Nova Sonic, Gemini Live, or OpenAI Realtime, that holds one connection open for the
whole conversation and listens while it responds. In your code, a model provider class
stands in for that service. It opens the connection and translates the service’s
streaming protocol into Strands events, so tools, hooks, I/O streams, and event
handling work the same whichever service you choose.
Realtime models are a separate family from the models a standard Agent uses, and
each provider class works only with BidiAgent.
Supported providers
Section titled “Supported providers”| Provider | Model class | Input | Connection limit |
|---|---|---|---|
| Amazon Nova Sonic | BedrockNovaSonicModel | Audio, text | 8 minutes |
| Google Gemini Live | GoogleGeminiLiveModel | Audio, text, images | About 10 minutes |
| OpenAI Realtime | OpenAIRealtimeModel | Audio, text, images | 60 minutes |
Each provider page covers credentials, client options, and behavior specific to that model.
Configuration
Section titled “Configuration”Every provider accepts the same core options, so moving between providers keeps the shape of your configuration:
| Option | Purpose |
|---|---|
model_id | The provider’s model identifier. Required. |
voice | The voice the model speaks with. Voice names are provider-specific. |
params | Provider-native session settings, such as temperature or turn detection. |
connection | Connection restart timing. See Connection limits. |
params is the pass-through to the provider’s own API. Field names and casing follow
that provider’s documentation, which each provider page links to.
from strands.bidi.models import BedrockNovaSonicModel
model = BedrockNovaSonicModel( model_id="amazon.nova-2-sonic-v1:0", voice="tiffany", params={"inferenceConfiguration": {"temperature": 0.7}},)To change configuration later, call update_config(). New values apply the next time
the connection opens, not to the connection already in progress.
Use the model’s audio configuration to set up capture and playback. The built-in
providers implement AudioCapable, which exposes
get_audio_config(). It returns separate input and output settings containing
each stream’s sample rate, channel count, and encoding.
AudioIO reads these settings automatically. Custom
I/O streams can read them to configure audio capture and
playback. The built-in providers use mono PCM; sample rates and configuration
options vary by provider. See each provider’s page for its supported settings.
Connection limits
Section titled “Connection limits”Every realtime provider limits how long a single connection stays open (see Supported providers). The agent restarts the connection shortly before that limit, at a break between turns when it can, and the conversation continues on the new connection. Providers carry context across the restart in one of two ways:
- History replay: Nova Sonic and OpenAI Realtime open a fresh connection and receive the conversation history.
- Session resumption: Gemini Live reconnects to the same server-side session through session resumption, falling back to history replay if resumption is unavailable.
Each provider sets its own default restart timing, and the connection option
overrides it or turns automatic restarts off. For the restart sequence and the events
it emits, see Connection restarts.
Custom providers
Section titled “Custom providers”To connect a realtime model Strands doesn’t support yet, subclass BidiModel. A
provider opens and closes the connection, sends user input and tool results, and
yields Strands events as the model responds, in the order described in
Event ordering. See the
BidiModel API reference for the
contract, and the built-in providers in strands/bidi/models/ for working examples.