Skip to content

Build a voice agent

Build an agent that listens and talks in real time. This guide walks through a bidirectional streaming agent end to end: audio input and output, streaming events, tool calls mid-conversation, and the model providers that support it.

After completing this guide, you can build voice assistants, interactive chatbots, multi-modal applications, and integrate bidirectional streaming with web servers or custom I/O streams.

Before starting, ensure you have:

  • Python 3.10+ installed (3.12+ required for Nova Sonic)
  • Audio hardware (microphone and speakers) for voice conversations
  • Model provider credentials configured (AWS, OpenAI, or Google)

Install the SDK with bidirectional streaming support:

To install support for all bidirectional streaming providers:

Terminal window
pip install "strands-agents[bidi-all]"

This includes all three providers (Nova Sonic, OpenAI, and Gemini Live), ConsoleIO, and microphone audio processing. Local microphone and speaker I/O with AudioIO also requires PortAudio and the bidi-pyaudio extra. See Platform-Specific Audio Setup.

You can also install support for specific providers:

Terminal window
# With local microphone and speaker I/O
pip install "strands-agents[bidi,bidi-io,bidi-pyaudio]"
# With terminal text I/O
pip install "strands-agents[bidi,bidi-io]"

AudioIO depends on PyAudio, which requires the PortAudio system library. Install PortAudio first, then install the bidi-pyaudio extra alongside bidi-all.

Terminal window
brew install portaudio
pip install "strands-agents[bidi-all,bidi-pyaudio]"

Bidirectional streaming supports multiple model providers. Choose one based on your needs:

Nova Sonic is Amazon’s bidirectional streaming model. Configure AWS credentials:

Terminal window
export AWS_ACCESS_KEY_ID=your_access_key
export AWS_SECRET_ACCESS_KEY=your_secret_key
export AWS_DEFAULT_REGION=us-east-1

Enable Nova Sonic model access in the Amazon Bedrock console.

Now let’s create a simple voice-enabled agent that can have real-time conversations:

import asyncio
from strands.bidi.agent import BidiAgent
from strands.bidi.io import AudioIO
from strands.bidi.models import BedrockNovaSonicModel
# Create a bidirectional streaming model
model = BedrockNovaSonicModel(model_id="amazon.nova-2-sonic-v1:0")
# Create the agent
agent = BidiAgent(
model=model,
system_prompt="You are a helpful voice assistant. Keep responses concise and natural."
)
# Setup audio I/O for microphone and speakers
audio_io = AudioIO()
# Run the conversation
async def main():
await agent.run(
inputs=[audio_io.input()],
outputs=[audio_io.output()]
)
asyncio.run(main())

You now have a voice-enabled agent that can:

  • Listen to your voice through the microphone
  • Process speech in real time
  • Respond with natural voice output
  • Display live user and assistant transcripts
  • Handle barge-ins when you start speaking

AudioIO.output() displays user and assistant transcripts while audio plays through the speakers. User speech appears in shaded > blocks and assistant speech appears as plain text.

The run() method runs indefinitely by default. The simplest way to stop conversations is using Ctrl+C:

import asyncio
from strands.bidi.agent import BidiAgent
from strands.bidi.io import AudioIO
from strands.bidi.models import BedrockNovaSonicModel
async def main():
model = BedrockNovaSonicModel(model_id="amazon.nova-2-sonic-v1:0")
agent = BidiAgent(model=model)
audio_io = AudioIO()
try:
# Runs indefinitely until interrupted
await agent.run(
inputs=[audio_io.input()],
outputs=[audio_io.output()]
)
except asyncio.CancelledError:
print("\nConversation cancelled by user")
finally:
# stop() should only be called after run() exits
await agent.stop()
asyncio.run(main())

Just like standard Strands agents, bidirectional agents can use tools during conversations:

import asyncio
from strands import tool
from strands.bidi.agent import BidiAgent
from strands.bidi.io import AudioIO
from strands.bidi.models import BedrockNovaSonicModel
from strands.vended_tools import notebook
# Define a custom tool
@tool
def get_weather(location: str) -> str:
"""
Get the current weather for a location.
Args:
location: City name or location
Returns:
Weather information
"""
# In a real application, call a weather API
return f"The weather in {location} is sunny and 72°F"
# Create agent with tools
model = BedrockNovaSonicModel(model_id="amazon.nova-2-sonic-v1:0")
agent = BidiAgent(
model=model,
tools=[notebook, get_weather],
system_prompt="You are a helpful assistant with access to tools."
)
audio_io = AudioIO()
async def main():
await agent.run(
inputs=[audio_io.input()],
outputs=[audio_io.output()]
)
asyncio.run(main())

You can now ask questions like:

  • “What time is it?”
  • “Calculate 25 times 48”
  • “What’s the weather in San Francisco?”

The agent automatically determines when to use tools and executes them concurrently without blocking the conversation.

Strands supports three bidirectional streaming providers:

  • Nova Sonic - Amazon’s bidirectional streaming model via AWS Bedrock
  • OpenAI Realtime - OpenAI’s Realtime API for voice conversations\
  • Gemini Live - Google’s multimodal streaming API

Each provider has different features, timeout limits, and audio quality. See the individual provider documentation for detailed configuration options.

Choose supported audio settings on the model and device buffering on the I/O stream:

import asyncio
from strands.bidi.agent import BidiAgent
from strands.bidi.io import AudioIO
from strands.bidi.models import GoogleGeminiLiveModel
# Configure model audio settings
model = GoogleGeminiLiveModel(
model_id="gemini-3.8-live",
audio={"input": {"sample_rate": 48000}},
voice="Puck",
)
# Configure I/O buffer settings
audio_io = AudioIO(
input_buffer_size=10, # Max input queue size
output_buffer_size=20, # Max output queue size
input_frames_per_buffer=512, # Input chunk size
output_frames_per_buffer=512 # Output chunk size
)
agent = BidiAgent(model=model)
async def main():
await agent.run(
inputs=[audio_io.input()],
outputs=[audio_io.output()]
)
asyncio.run(main())

AudioIO reads the model’s resolved input and output formats through get_audio_config(). You do not need to repeat rates or channel counts on the I/O stream.

Bidirectional agents automatically handle barge-ins when users start speaking:

import asyncio
from strands.bidi.agent import BidiAgent
from strands.bidi.io import AudioIO
from strands.bidi.models import BedrockNovaSonicModel
from strands.bidi.types import BidiBargeInEvent
model = BedrockNovaSonicModel(model_id="amazon.nova-2-sonic-v1:0")
agent = BidiAgent(model=model)
audio_io = AudioIO()
async def main():
await agent.start()
# Start receiving events
async for event in agent.receive():
if isinstance(event, BidiBargeInEvent):
print(f"Barge-in: {event.reason}")
# Audio output automatically cleared
# Model stops generating
# Ready for new input
asyncio.run(main())

Barge-ins are detected via voice activity detection (VAD) and handled automatically:

  1. User starts speaking
  2. Model stops generating
  3. Audio output buffer cleared
  4. Model ready for new input

If you need more control over the agent lifecycle, you can manually call start() and stop():

import asyncio
from strands.bidi.agent import BidiAgent
from strands.bidi.models import BedrockNovaSonicModel
from strands.bidi.types import BidiResponseStopEvent
async def main():
model = BedrockNovaSonicModel(model_id="amazon.nova-2-sonic-v1:0")
agent = BidiAgent(model=model)
# Manually start the agent
await agent.start()
try:
await agent.send("What is Python?")
async for event in agent.receive():
if isinstance(event, BidiResponseStopEvent):
break
finally:
# Always stop after exiting receive loop
await agent.stop()
asyncio.run(main())

See Controlling Conversation Lifecycle for more patterns and best practices.

To let users end a conversation by voice, define a tool that calls agent.cancel():

import asyncio
from strands import LocalAgent, ToolContext, tool
from strands.bidi.agent import BidiAgent
from strands.bidi.io import AudioIO
from strands.bidi.models import BedrockNovaSonicModel
@tool(context=True)
def end_conversation(tool_context: ToolContext[LocalAgent]) -> str:
"""End the conversation when the user asks to stop."""
tool_context.agent.cancel()
return "Ending conversation"
model = BedrockNovaSonicModel(model_id="amazon.nova-2-sonic-v1:0")
agent = BidiAgent(
model=model,
tools=[end_conversation],
system_prompt="You are a helpful assistant.",
)
audio_io = AudioIO()
async def main():
await agent.run(
inputs=[audio_io.input()],
outputs=[audio_io.output()]
)
# run() returns after the agent calls end_conversation.
asyncio.run(main())

Bidi checks for cancellation after a tool group completes and its results are recorded. Requests from other contexts remain pending until that checkpoint.

To enable debug logs in your agent, configure the strands logger:

import asyncio
import logging
from strands.bidi.agent import BidiAgent
from strands.bidi.io import AudioIO
from strands.bidi.models import BedrockNovaSonicModel
# Enable debug logs
logging.getLogger("strands").setLevel(logging.DEBUG)
logging.basicConfig(
format="%(levelname)s | %(name)s | %(message)s",
handlers=[logging.StreamHandler()]
)
model = BedrockNovaSonicModel(model_id="amazon.nova-2-sonic-v1:0")
agent = BidiAgent(model=model)
audio_io = AudioIO()
async def main():
await agent.run(
inputs=[audio_io.input()],
outputs=[audio_io.output()]
)
asyncio.run(main())

Debug logs show:

  • Connection lifecycle events
  • Audio buffer operations
  • Tool execution details
  • Event processing flow

Over open speakers, the agent’s own playback can feed back into the microphone and trigger a barge-in. Either use a headset, or enable microphone audio processing to cancel the echo:

Terminal window
pip install "strands-agents[bidi,bidi-pyaudio,bidi-aec]"
audio_io = AudioIO(audio_processor=True)

See Audio Processing for the available options.

If you don’t hear audio:

# List available audio devices
import pyaudio
p = pyaudio.PyAudio()
for i in range(p.get_device_count()):
info = p.get_device_info_by_index(i)
print(f"{i}: {info['name']}")
# Specify output device explicitly
audio_io = AudioIO(output_device_index=2)

If the agent doesn’t respond to speech:

# Specify input device explicitly
audio_io = AudioIO(input_device_index=1)
# Check system permissions (macOS)
# System Preferences → Security & Privacy → Microphone

Each provider caps how long a single connection stays open. Rather than wait for that limit, BidiAgent restarts the connection proactively: a timer fires ahead of the cap, the agent replays the conversation history into a fresh connection, and it emits a BidiConnectionRestartEvent with reason="scheduled". If a connection times out first, the agent restarts reactively and emits the same event with reason="timeout". Treat both as informational, not errors:

from strands.bidi.types import BidiConnectionRestartEvent
async for event in agent.receive():
if isinstance(event, BidiConnectionRestartEvent):
print(f"Restarting (reason={event.reason})")
if event.turn_interrupted:
print("The in-progress turn was cut short; consider re-prompting.")
continue

Providers declare restart timing through ConnectionConfig. Tune it, or opt out of automatic restart with the model’s connection argument:

from strands.bidi.models import BedrockNovaSonicModel
# Restart 60s earlier than the provider default
model = BedrockNovaSonicModel(
model_id="amazon.nova-2-sonic-v1:0",
connection={"restart_after_s": 360},
)

For longer sessions on a single connection, OpenAI Realtime allows a larger connection window than Nova Sonic.

  • Agent - Deep dive into BidiAgent configuration and lifecycle
  • Events - Complete guide to bidirectional streaming events
  • I/O Streams - Understanding and customizing input and output streams
  • Model Providers:
  • Python API Reference - Complete API documentation