Build a voice agent
Build an agent that listens and talks in real time. This guide walks through a bidirectional streaming agent end to end: audio input and output, streaming events, tool calls mid-conversation, and the model providers that support it.
After completing this guide, you can build voice assistants, interactive chatbots, multi-modal applications, and integrate bidirectional streaming with web servers or custom I/O channels.
Prerequisites
Section titled “Prerequisites”Before starting, ensure you have:
- Python 3.10+ installed (3.12+ required for Nova Sonic)
- Audio hardware (microphone and speakers) for voice conversations
- Model provider credentials configured (AWS, OpenAI, or Google)
Install the SDK
Section titled “Install the SDK”Bidirectional streaming is included in the Strands Agents SDK as an experimental feature. Install the SDK with bidirectional streaming support:
For All Providers
Section titled “For All Providers”To install support for all bidirectional streaming providers:
pip install "strands-agents[bidi-all]"This includes all three providers (Nova Sonic, OpenAI, and Gemini Live), BidiTextIO, and microphone audio processing. For local microphone and speaker I/O with BidiAudioIO, also install PortAudio and the bidi-pyaudio extra (see Platform-Specific Audio Setup); PyAudio is excluded from bidi-all because of its PortAudio system dependency.
For Specific Providers
Section titled “For Specific Providers”You can also install support for specific providers:
# With local microphone and speaker I/Opip install "strands-agents[bidi,bidi-io,bidi-pyaudio]"
# With terminal text I/Opip install "strands-agents[bidi,bidi-io]"# With local audio I/Opip install "strands-agents[bidi-io,bidi-openai,bidi-pyaudio]"
# With terminal text I/Opip install "strands-agents[bidi-io,bidi-openai]"# With local audio I/Opip install "strands-agents[bidi-google,bidi-io,bidi-pyaudio]"
# With terminal text I/Opip install "strands-agents[bidi-google,bidi-io]"Platform-Specific Audio Setup
Section titled “Platform-Specific Audio Setup”BidiAudioIO depends on PyAudio, which requires the PortAudio system library. Install PortAudio first, then install the bidi-pyaudio extra alongside bidi-all.
brew install portaudiopip install "strands-agents[bidi-all,bidi-pyaudio]"sudo apt-get install portaudio19-dev python3-pyaudiopip install "strands-agents[bidi-all,bidi-pyaudio]"PyAudio typically installs without additional dependencies.
pip install "strands-agents[bidi-all,bidi-pyaudio]"Configuring Credentials
Section titled “Configuring Credentials”Bidirectional streaming supports multiple model providers. Choose one based on your needs:
Nova Sonic is Amazon’s bidirectional streaming model. Configure AWS credentials:
export AWS_ACCESS_KEY_ID=your_access_keyexport AWS_SECRET_ACCESS_KEY=your_secret_keyexport AWS_DEFAULT_REGION=us-east-1Enable Nova Sonic model access in the Amazon Bedrock console.
For OpenAI’s Realtime API, set your API key:
export OPENAI_API_KEY=your_api_keyFor Gemini Live API, set your API key:
export GOOGLE_API_KEY=your_api_keyYour First Voice Conversation
Section titled “Your First Voice Conversation”Now let’s create a simple voice-enabled agent that can have real-time conversations:
import asynciofrom strands.experimental.bidi.agent import BidiAgentfrom strands.experimental.bidi.io import BidiAudioIOfrom strands.experimental.bidi.models import BedrockNovaSonicModel
# Create a bidirectional streaming modelmodel = BedrockNovaSonicModel()
# Create the agentagent = BidiAgent( model=model, system_prompt="You are a helpful voice assistant. Keep responses concise and natural.")
# Setup audio I/O for microphone and speakersaudio_io = BidiAudioIO()
# Run the conversationasync def main(): await agent.run( inputs=[audio_io.input()], outputs=[audio_io.output()] )
asyncio.run(main())You now have a voice-enabled agent that can:
- Listen to your voice through the microphone
- Process speech in real time
- Respond with natural voice output
- Display live user and assistant transcripts
- Handle interruptions when you start speaking
Live Transcripts
Section titled “Live Transcripts”BidiAudioIO.output() displays user and assistant transcripts while audio plays
through the speakers. User speech appears in shaded > blocks and assistant
speech appears as plain text.
Controlling Conversation Lifecycle
Section titled “Controlling Conversation Lifecycle”The run() method runs indefinitely by default. The simplest way to stop conversations is using Ctrl+C:
import asynciofrom strands.experimental.bidi.agent import BidiAgentfrom strands.experimental.bidi.io import BidiAudioIOfrom strands.experimental.bidi.models import BedrockNovaSonicModel
async def main(): model = BedrockNovaSonicModel() agent = BidiAgent(model=model) audio_io = BidiAudioIO()
try: # Runs indefinitely until interrupted await agent.run( inputs=[audio_io.input()], outputs=[audio_io.output()] ) except asyncio.CancelledError: print("\nConversation cancelled by user") finally: # stop() should only be called after run() exits await agent.stop()
asyncio.run(main())Adding Tools to Your Agent
Section titled “Adding Tools to Your Agent”Just like standard Strands agents, bidirectional agents can use tools during conversations:
import asynciofrom strands import toolfrom strands.experimental.bidi.agent import BidiAgentfrom strands.experimental.bidi.io import BidiAudioIOfrom strands.experimental.bidi.models import BedrockNovaSonicModelfrom strands.vended_tools import notebook
# Define a custom tool@tooldef get_weather(location: str) -> str: """ Get the current weather for a location.
Args: location: City name or location
Returns: Weather information """ # In a real application, call a weather API return f"The weather in {location} is sunny and 72°F"
# Create agent with toolsmodel = BedrockNovaSonicModel()agent = BidiAgent( model=model, tools=[notebook, get_weather], system_prompt="You are a helpful assistant with access to tools.")
audio_io = BidiAudioIO()
async def main(): await agent.run( inputs=[audio_io.input()], outputs=[audio_io.output()] )
asyncio.run(main())You can now ask questions like:
- “What time is it?”
- “Calculate 25 times 48”
- “What’s the weather in San Francisco?”
The agent automatically determines when to use tools and executes them concurrently without blocking the conversation.
Model Providers
Section titled “Model Providers”Strands supports three bidirectional streaming providers:
- Nova Sonic - Amazon’s bidirectional streaming model via AWS Bedrock
- OpenAI Realtime - OpenAI’s Realtime API for voice conversations\
- Gemini Live - Google’s multimodal streaming API
Each provider has different features, timeout limits, and audio quality. See the individual provider documentation for detailed configuration options.
Configuring Audio Settings
Section titled “Configuring Audio Settings”Choose supported audio settings on the model and device buffering on the I/O channel:
import asyncio
from strands.experimental.bidi.agent import BidiAgentfrom strands.experimental.bidi.io import BidiAudioIOfrom strands.experimental.bidi.models import GoogleGeminiLiveModel
# Configure model audio settingsmodel = GoogleGeminiLiveModel( audio={"input": {"sample_rate": 48000}}, voice="Puck",)
# Configure I/O buffer settingsaudio_io = BidiAudioIO( input_buffer_size=10, # Max input queue size output_buffer_size=20, # Max output queue size input_frames_per_buffer=512, # Input chunk size output_frames_per_buffer=512 # Output chunk size)
agent = BidiAgent(model=model)
async def main(): await agent.run( inputs=[audio_io.input()], outputs=[audio_io.output()] )
asyncio.run(main())BidiAudioIO reads the model’s resolved input and output formats through
get_audio_config(). You do not need to repeat rates or channel counts on the I/O
channel.
Handling Interruptions
Section titled “Handling Interruptions”Bidirectional agents automatically handle interruptions when users start speaking:
import asynciofrom strands.experimental.bidi.agent import BidiAgentfrom strands.experimental.bidi.io import BidiAudioIOfrom strands.experimental.bidi.models import BedrockNovaSonicModelfrom strands.experimental.bidi.types import BidiInterruptionEvent
model = BedrockNovaSonicModel()agent = BidiAgent(model=model)audio_io = BidiAudioIO()
async def main(): await agent.start()
# Start receiving events async for event in agent.receive(): if isinstance(event, BidiInterruptionEvent): print(f"User interrupted: {event.reason}") # Audio output automatically cleared # Model stops generating # Ready for new input
asyncio.run(main())Interruptions are detected via voice activity detection (VAD) and handled automatically:
- User starts speaking
- Model stops generating
- Audio output buffer cleared
- Model ready for new input
Manual Start and Stop
Section titled “Manual Start and Stop”If you need more control over the agent lifecycle, you can manually call start() and stop():
import asynciofrom strands.experimental.bidi.agent import BidiAgentfrom strands.experimental.bidi.models import BedrockNovaSonicModelfrom strands.experimental.bidi.types import BidiResponseCompleteEvent
async def main(): model = BedrockNovaSonicModel() agent = BidiAgent(model=model)
# Manually start the agent await agent.start()
try: await agent.send("What is Python?")
async for event in agent.receive(): if isinstance(event, BidiResponseCompleteEvent): break finally: # Always stop after exiting receive loop await agent.stop()
asyncio.run(main())See Controlling Conversation Lifecycle for more patterns and best practices.
Graceful Shutdown
Section titled “Graceful Shutdown”Use the SDK’s experimental stop tool to allow users to end conversations
naturally. It sets request_state["stop_event_loop"], which the agent loop checks
to trigger a graceful shutdown:
import asynciofrom strands.experimental.bidi.agent import BidiAgentfrom strands.experimental.bidi.io import BidiAudioIOfrom strands.experimental.bidi.models import BedrockNovaSonicModelfrom strands.experimental.tools import stop
model = BedrockNovaSonicModel()agent = BidiAgent( model=model, tools=[stop], system_prompt="You are a helpful assistant. When the user says 'stop conversation', use the stop tool.")
audio_io = BidiAudioIO()
async def main(): await agent.run( inputs=[audio_io.input()], outputs=[audio_io.output()] ) # Conversation ends when user says "stop conversation"
asyncio.run(main())You can also create custom stop tools using the request_state["stop_event_loop"] flag:
from strands import tool
@tooldef end_session(request_state: dict) -> str: request_state["stop_event_loop"] = True return "Goodbye!"The agent will gracefully close the connection when any tool sets request_state["stop_event_loop"] = True.
Debug Logs
Section titled “Debug Logs”To enable debug logs in your agent, configure the strands logger:
import asyncioimport loggingfrom strands.experimental.bidi.agent import BidiAgentfrom strands.experimental.bidi.io import BidiAudioIOfrom strands.experimental.bidi.models import BedrockNovaSonicModel
# Enable debug logslogging.getLogger("strands").setLevel(logging.DEBUG)logging.basicConfig( format="%(levelname)s | %(name)s | %(message)s", handlers=[logging.StreamHandler()])
model = BedrockNovaSonicModel()agent = BidiAgent(model=model)audio_io = BidiAudioIO()
async def main(): await agent.run( inputs=[audio_io.input()], outputs=[audio_io.output()] )
asyncio.run(main())Debug logs show:
- Connection lifecycle events
- Audio buffer operations
- Tool execution details
- Event processing flow
Common Issues
Section titled “Common Issues”Audio Feedback Loop in a Python Console
Section titled “Audio Feedback Loop in a Python Console”Over open speakers, the agent’s own playback can feed back into the microphone and interrupt it. Either use a headset, or enable microphone audio processing to cancel the echo:
pip install "strands-agents[bidi,bidi-pyaudio,bidi-aec]"audio_io = BidiAudioIO(audio_processor=True)See Audio Processing for the available options.
No Audio Output
Section titled “No Audio Output”If you don’t hear audio:
# List available audio devicesimport pyaudiop = pyaudio.PyAudio()for i in range(p.get_device_count()): info = p.get_device_info_by_index(i) print(f"{i}: {info['name']}")
# Specify output device explicitlyaudio_io = BidiAudioIO(output_device_index=2)Microphone Not Working
Section titled “Microphone Not Working”If the agent doesn’t respond to speech:
# Specify input device explicitlyaudio_io = BidiAudioIO(input_device_index=1)
# Check system permissions (macOS)# System Preferences → Security & Privacy → MicrophoneConnection Restarts
Section titled “Connection Restarts”Each provider caps how long a single connection stays open. Rather than wait for that limit, BidiAgent reconnects proactively: a timer fires ahead of the cap, the agent replays the conversation history into a fresh connection, and it emits a BidiConnectionRestartEvent with reason="scheduled". If a connection times out first, the agent reconnects reactively and emits the same event with reason="timeout". Treat both as informational, not errors:
from strands.experimental.bidi.types import BidiConnectionRestartEvent
async for event in agent.receive(): if isinstance(event, BidiConnectionRestartEvent): print(f"Reconnecting (reason={event.reason})") if event.turn_interrupted: print("The in-progress turn was cut short; consider re-prompting.") continueProviders declare the reconnect timing through BidiConnectionConfig. Tune it, or opt out of automatic reconnect, through provider_config["connection"]:
from strands.experimental.bidi.models import BedrockNovaSonicModel
# Reconnect 60s earlier than the provider defaultmodel = BedrockNovaSonicModel(provider_config={"connection": {"restart_after_s": 360}})For longer sessions on a single connection, OpenAI Realtime allows a larger connection window than Nova Sonic.
Next Steps
Section titled “Next Steps”- Agent - Deep dive into BidiAgent configuration and lifecycle
- Events - Complete guide to bidirectional streaming events
- I/O Channels - Understanding and customizing input/output channels
- Model Providers:
- Nova Sonic - Amazon Bedrock’s bidirectional streaming model
- OpenAI Realtime - OpenAI’s Realtime API
- Gemini Live - Google’s Gemini Live API
- Python API Reference - Complete API documentation