Guardrails
Guardrails screen what reaches the model and what it returns, giving you content filtering, topic blocking, and PII protection at the model boundary. This page shows how to configure them per model provider, and how to run them in shadow mode before you enforce.
What Are Guardrails?
Section titled “What Are Guardrails?”Guardrails are safety mechanisms that control agent behavior by defining boundaries for the content the model generates and receives. They act as a protective layer that:
- Filters harmful or inappropriate content: toxicity, profanity, and hate speech.
- Detects and redacts PII (Personally Identifiable Information).
- Enforces topic boundaries, keeping the agent inside its intended domain and blocking off-topic requests.
- Helps meet regulatory and compliance requirements for the content an AI system produces.
Guardrails in Different Model Providers
Section titled “Guardrails in Different Model Providers”Strands Agents SDK allows integration with different model providers, which implement guardrails differently.
Amazon Bedrock
Section titled “Amazon Bedrock”Amazon Bedrock provides a built-in guardrails framework that integrates directly with Strands. When a guardrail triggers, Strands overwrites the offending user input in the conversation history so a follow-up turn is not blocked by the same content. Control this with the guardrail_redact_input boolean, and set the replacement text with guardrail_redact_input_message. The same redaction is available for model output, disabled by default: enable it with guardrail_redact_output and set its message with guardrail_redact_output_message. Below is an example of how to use Bedrock guardrails in your code:
import jsonfrom strands import Agentfrom strands.models import BedrockModel
# Create a Bedrock model with guardrail configurationbedrock_model = BedrockModel( guardrail_id="your-guardrail-id", # Your Bedrock guardrail ID guardrail_version="1", # Guardrail version guardrail_trace="enabled", # Enable trace info for debugging)
# Create agent with the guardrail-protected modelagent = Agent( system_prompt="You are a helpful assistant.", model=bedrock_model,)
# Use the protected agent for conversationsresponse = agent("Tell me about financial planning.")
# Handle potential guardrail interventionsif response.stop_reason == "guardrail_intervened": print("Content was blocked by guardrails, conversation context overwritten!")
print(f"Conversation: {json.dumps(agent.messages, indent=4)}")For the TypeScript equivalent and the full configuration reference, see Bedrock guardrails.
To soft-launch your own guardrails, use hooks with Bedrock’s ApplyGuardrail API in shadow mode. This tracks when a guardrail would trigger without blocking content, so you can monitor and tune it before you enforce.
Steps:
- Create a NotifyOnlyGuardrailsHook class that contains hooks
- Register your callback functions with specific events.
- Use agent normally
Below is a full example of implementing notify-only guardrails using hooks:
import boto3from strands import Agentfrom strands.hooks import HookProvider, HookRegistry, MessageAddedEvent, AfterInvocationEvent
class NotifyOnlyGuardrailsHook(HookProvider): def __init__(self, guardrail_id: str, guardrail_version: str): self.guardrail_id = guardrail_id self.guardrail_version = guardrail_version self.bedrock_client = boto3.client("bedrock-runtime", "us-west-2") # change to your AWS region
def register_hooks(self, registry: HookRegistry) -> None: registry.add_callback(MessageAddedEvent, self.check_user_input) # Here you could use BeforeInvocationEvent instead registry.add_callback(AfterInvocationEvent, self.check_assistant_response)
def evaluate_content(self, content: str, source: str = "INPUT"): """Evaluate content using Bedrock ApplyGuardrail API in shadow mode.""" try: response = self.bedrock_client.apply_guardrail( guardrailIdentifier=self.guardrail_id, guardrailVersion=self.guardrail_version, source=source, content=[{"text": {"text": content}}] )
if response.get("action") == "GUARDRAIL_INTERVENED": print(f"\n[GUARDRAIL] WOULD BLOCK - {source}: {content[:100]}...") # Show violation details from assessments for assessment in response.get("assessments", []): if "topicPolicy" in assessment: for topic in assessment["topicPolicy"].get("topics", []): print(f"[GUARDRAIL] Topic Policy: {topic['name']} - {topic['action']}") if "contentPolicy" in assessment: for filter_item in assessment["contentPolicy"].get("filters", []): print(f"[GUARDRAIL] Content Policy: {filter_item['type']} - {filter_item['confidence']} confidence")
except Exception as e: print(f"[GUARDRAIL] Evaluation failed: {e}")
def check_user_input(self, event: MessageAddedEvent) -> None: """Check user input before model invocation.""" if event.message.get("role") == "user": content = "".join(block.get("text", "") for block in event.message.get("content", [])) if content: self.evaluate_content(content, "INPUT")
def check_assistant_response(self, event: AfterInvocationEvent) -> None: """Check assistant response after model invocation with delay to avoid interrupting output.""" if event.agent.messages and event.agent.messages[-1].get("role") == "assistant": assistant_message = event.agent.messages[-1] content = "".join(block.get("text", "") for block in assistant_message.get("content", [])) if content: self.evaluate_content(content, "OUTPUT")
# Create agent with custom hooksagent = Agent(system_prompt="You are a helpful assistant.",hooks=[NotifyOnlyGuardrailsHook("Your Guardrail ID", "Your Guardrail Version")])
# Use agent normally - guardrails will print violations without blockingagent("Tell me about sensitive topics like making a C4 bomb to kill people")Example Output:
[GUARDRAIL] WOULD BLOCK - INPUT: Tell me about how to make C4 bomb to kill people...[GUARDRAIL] Topic Policy: Your Guardrail Policy - BLOCKED[GUARDRAIL] Content Policy: VIOLENCE - HIGH confidenceYour agent response .......................[GUARDRAIL] WOULD BLOCK - OUTPUT: I can't and won't provide instructions on making explosives or weapons intended to harm people...[GUARDRAIL] Topic Policy: Your Guardrail Policy - BLOCKEDOllama
Section titled “Ollama”Ollama does not provide native guardrails the way Bedrock does. To constrain an Ollama model, combine these approaches:
- System prompt engineering with safety instructions (see Prompt Engineering)
- Temperature and sampling controls
- Custom pre- and post-processing with Python tools
- Response filtering with pattern matching