Skip to content

Anthropic

Anthropic is an AI safety and research company and the maker of the Claude family of models. The Strands Agents SDK implements an Anthropic provider, letting you run agents against Claude models directly.

Anthropic is configured as an optional dependency in Strands Agents. To install, run:

Terminal window
pip install 'strands-agents[anthropic]'

After installing dependencies, you can import and initialize the Strands Agents’ Anthropic provider as follows:

from strands import Agent
from strands.models.anthropic import AnthropicModel
from strands.vended_tools import notebook
model = AnthropicModel(
client_args={
"api_key": "<KEY>",
},
# **model_config
max_tokens=1028,
model_id="claude-sonnet-5",
params={
"temperature": 0.7,
}
)
agent = Agent(model=model, tools=[notebook])
response = agent('Create a notebook named "ideas" and add three project ideas.')
print(response)

The client_args configure the underlying Anthropic client. For a complete list of available arguments, please refer to the Anthropic Python SDK docs.

The model_config configures the underlying model selected for inference. The supported configurations are:

ParameterDescriptionExampleOptions
max_tokensMaximum number of tokens to generate before stopping1028reference
model_idID of a model to useclaude-sonnet-5reference
paramsAdditional pass-through parameters{"metadata": {"user_id": "u1"}}reference
anthropic_toolsBuilt-in tools, appended to the agent’s function tools[{"type": "web_search_20260318", "name": "web_search"}]reference
cache_configEnables prompt caching on the system prompt and the conversationCacheConfig(ttl="1h")reference
cache_toolsCaches the tool definitions (deprecated, use cache_config with tools_ttl)"default" or CacheToolsConfig(ttl="1h")reference

If you encounter the error ModuleNotFoundError: No module named 'anthropic', this means you haven’t installed the anthropic dependency in your environment. To fix, run pip install 'strands-agents[anthropic]'.

You can pass a pre-configured Anthropic client directly to AnthropicModel. You are responsible for managing the client’s lifecycle.

The Python SDK does not currently support passing a pre-configured client. Use client_args to configure the client at initialization.

Anthropic models support structured output through tool use. Pass a schema to the agent, and Strands generates a tool from it that the model calls to return validated, type-safe data.

Define a Pydantic model and pass it to agent.structured_output():

from pydantic import BaseModel, Field
from strands import Agent
from strands.models.anthropic import AnthropicModel
class MovieReview(BaseModel):
"""Analyze a movie review."""
title: str = Field(description="Movie title")
rating: int = Field(description="Rating from 1-10", ge=1, le=10)
genre: str = Field(description="Primary genre")
sentiment: str = Field(description="Overall sentiment: positive, negative, or neutral")
summary: str = Field(description="Brief summary of the review")
model = AnthropicModel(
client_args={"api_key": "<KEY>"},
max_tokens=1028,
model_id="claude-sonnet-5",
)
agent = Agent(model=model)
result = agent.structured_output(
MovieReview,
"""
Just watched "The Matrix" - what an incredible sci-fi masterpiece!
The groundbreaking visual effects and philosophical themes make this
a must-watch. Keanu Reeves delivers a solid performance. 9/10!
"""
)
print(f"Movie: {result.title}")
print(f"Rating: {result.rating}/10")
print(f"Genre: {result.genre}")
print(f"Sentiment: {result.sentiment}")

For schema patterns, error handling, and per-invocation overrides, see Structured Output.

Anthropic’s built-in server-side tools (web search, web fetch, code execution) can be passed via the anthropic_tools config option. These are appended alongside any function tools registered on the agent.

from strands import Agent
from strands.models.anthropic import AnthropicModel
model = AnthropicModel(
client_args={"api_key": "<KEY>"},
model_id="claude-sonnet-4-6",
max_tokens=1028,
anthropic_tools=[{"type": "web_search_20260318", "name": "web_search", "max_uses": 5}],
)
agent = Agent(model=model)
response = agent("What are the latest AI news today?")
print(response)

Web search results are surfaced as citations on the response text when Claude calls the tool directly (allowed_callers: ["direct"]). On web_search_20260318 the default is dynamic filtering, which runs the search inside code execution and returns no citations. For available built-in tools and their versioned type strings, see the Anthropic tool use documentation.

Limitations:

  • The raw server-tool blocks (search results, fetched pages, code output) are not kept in the conversation history, so on a later turn the model cannot refer back to them and will call the tool again if it needs them.
  • When Anthropic pauses a long-running server-tool turn, the model provider resumes it automatically, up to 10 times per request, before the response reaches the agent loop. If the turn is still paused after that, the provider raises an error.
  • Server tools are not sent when a specific tool call is forced (tool_choice of any or tool), which includes the forced structured-output retry; the model can only use them on turns where it is free to choose.
  • Server-tool calls are billed separately by Anthropic and are not included in Strands usage metrics.

Prompt caching lets Claude reuse an already-processed prefix of your prompt instead of reprocessing it on every call. The mechanism, the cost model, and the cache metrics match Amazon Bedrock; this section covers what differs here.

Caching is off by default. Set cache_configcacheConfig to cache the system prompt and add a cache point to the last user message, which caches everything before it:

from strands import Agent
from strands.models.anthropic import AnthropicModel
from strands.models import CacheConfig
model = AnthropicModel(
model_id="claude-sonnet-5",
max_tokens=1028,
cache_config=CacheConfig(ttl="1h", tools_ttl="1h"),
)
agent = Agent(model=model)
result = agent("Summarize the attached report.")
print(result.metrics.accumulated_usage)
# Typical output:
# {'inputTokens': 12, ..., 'cacheReadInputTokens': 2505}

Cache activity is reported in the usage metrics, in the same fields Bedrock uses: see cache metrics.

cache_config caches the system prompt and the conversation; set system_prompt_ttl=False to leave the system prompt uncached, or a TTL string (system_prompt_ttl="1h") to give it its own duration. Add tools_ttl to the same cache_config to cache the tool definitions as well: a TTL string sets the duration, True uses the provider default. The model-level cache_tools parameter (and CacheToolsConfig) is deprecated in favor of tools_ttl; it still works, and unlike tools_ttl it caches the tool definitions on its own, without a cache_config. Passing a plain string ("default") to cache_tools only switches it on: the value is not a TTL.

strategy is ignored by this provider.

Placement works as it does on Amazon Bedrock: a cache point you put in the last user message is kept where you put it, automatic placement is suspended for that message, and points in earlier messages are removed. Place your point ahead of content that is rebuilt on every call, otherwise it lands inside the cached prefix and every request writes an entry that none ever reads.

Two constraints are specific to this API:

  • Only some block types accept a cache point. Text, image, tool use, tool result, and document blocks do; a reasoning block does not. A point with only a reasoning block, or a media block sourced by location, ahead of it cannot be honored, so automatic placement applies instead.
  • Maximum of four cache points per request, shared across the tool definitions, the system prompt, and the messages. Automatic placement uses one per part it caches, so hand-placing several of your own alongside it can exceed the limit, and the API rejects the request.

A prompt below the model’s minimum cacheable length is not cached: if both cache metrics stay at zero, that is the likely cause. See cache limitations for per-model thresholds.

Token counting is used by context management strategies to estimate input tokens before each model call.

The Anthropic provider can use the native messages.count_tokens() API, which provides exact token counts including system prompts, messages, and tool specifications.

You can enable native token counting with:

model = AnthropicModel(
model_id="claude-sonnet-5",
use_native_token_count=True,
)

When disabled (or if the API call fails), falls back to estimation with a character-based heuristic (characters ÷ 4 for text, characters ÷ 2 for JSON).