Agent configuration
Agent Configuration
Configure the agent with the fluent builder. Use OpenAIAgent.builder[F] / ClaudeAgent.builder[F] for an effect system (e.g. builder[IO]), or synchronous(...) for the blocking Identity clients:
val agent = OpenAIAgent
.synchronous(OpenAI.fromEnv, "gpt-4o-mini")
.maxIterations(10) // Max reasoning steps
.maxTokens(8192) // Optional per-call output-token cap
.systemPrompt("Custom prompt") // Optional instructions
.tools(tool1, tool2) // Your tools
.deriveResponseSchema[T] // (fixes the agent's output type — see typed input and output in tools.md)
.build
maxIterations bounds the total number of model calls. The last allowed iteration is sent without tools, forcing the
model to produce a final text answer instead of a tool call whose result would be discarded (that iteration finishes with
FinishReason.MaxIterations). Tools are therefore only available for the first maxIterations - 1 iterations — so with
maxIterations = 1 the agent never gets to use its tools. Set maxIterations at least one higher than the number of
tool-using steps you expect the task to need.
maxTokens caps the tokens the model may generate on each LLM call. When unset, each provider’s default applies:
Claude and Gemini send 4096, OpenAI sends no cap (the model’s own maximum applies). For OpenAI it is sent as
max_completion_tokens, which also works with reasoning models. When a response is cut off by the cap, the run ends
with FinishReason.TokenLimit and the partial answer is returned as Left(AgentIncomplete(...)).
The OpenAI factories additionally accept a strictTools flag (default true): when true, tool schemas are
normalized for OpenAI’s strict function calling (additionalProperties: false, all properties required, optional
properties nullable); when false, tools are registered as non-strict with their original schemas.
Model capabilities
Model constants are tagged with capability marker traits: Vision, ToolCalling, StructuredOutput, Reasoning
(in sttp.ai.core.model.Capability; models supporting everything mix in the Capability.All shorthand, which
extends all four). Agent builder methods that need a capability require it at compile time:
.tools/.addTool need ToolCalling, .responseSchema/.deriveResponseSchema need StructuredOutput.
import sttp.ai.core.agent.AgentTool
import sttp.ai.openai.OpenAI
import sttp.ai.openai.agent.OpenAIAgent
import sttp.ai.openai.requests.completions.chat.ChatRequestBody.ChatCompletionModel
import sttp.shared.Identity
def myTool: AgentTool[Identity, ?] = ???
val agent = OpenAIAgent
.synchronous(OpenAI.fromEnv, ChatCompletionModel.GPT4o) // GPT4o mixes in ToolCalling
.tools(myTool)
.build
// OpenAIAgent.synchronous(OpenAI.fromEnv, ChatCompletionModel.O1Mini).tools(myTool) // does not compile: o1-mini has no ToolCalling
String model names keep working and skip capability checking — they wrap into the provider’s custom model class, which claims all capabilities (you assert your model supports what you use it for):
import sttp.ai.core.agent.AgentTool
import sttp.ai.openai.OpenAI
import sttp.ai.openai.agent.OpenAIAgent
import sttp.shared.Identity
def myTool: AgentTool[Identity, ?] = ???
val ollamaAgent = OpenAIAgent.synchronous(OpenAI.fromEnv, "llama3-70b").tools(myTool).build
Per-iteration model selection
Agents can use a different model per loop iteration — e.g. a cheap model while the agent calls tools, and a stronger model for the forced-final iteration (where tools are withheld and the model must answer):
import sttp.ai.core.agent.{AgentTool, IterationInfo}
import sttp.ai.openai.OpenAI
import sttp.ai.openai.agent.OpenAIAgent
import sttp.ai.openai.requests.completions.chat.ChatRequestBody.ChatCompletionModel
import sttp.shared.Identity
def myTool: AgentTool[Identity, ?] = ???
val mixedAgent = OpenAIAgent
.synchronous(OpenAI.fromEnv, (info: IterationInfo) => if info.isLastIteration then ChatCompletionModel.GPT5 else ChatCompletionModel.GPT4oMini)
.tools(myTool)
.build
The inferred model type is the least upper bound of every model the function can return, so capability checks require the capabilities all of them share. If inference produces an unwieldy type, ascribe the function explicitly:
import sttp.ai.core.agent.IterationInfo
import sttp.ai.core.model.Capability
import sttp.ai.openai.requests.completions.chat.ChatRequestBody.ChatCompletionModel
val pick: IterationInfo => ChatCompletionModel & Capability.ToolCalling =
info => if info.isLastIteration then ChatCompletionModel.GPT5 else ChatCompletionModel.GPT4oMini
Note: isLastIteration is true on the forced last iteration (maxIterations reached) and on an
interceptor-forced final iteration (e.g. a budget breach). The loop cannot know in
advance on which iteration the model will answer naturally.
Exception Handling
The ExceptionHandler controls how tool execution errors and argument parsing failures are handled. You can choose between built-in handlers or create custom ones.
Built-in Handlers:
Handler |
Tool Execution Errors |
Parse Errors |
Use Case |
|---|---|---|---|
|
IO/Interrupt errors propagate; others sent to LLM |
Sent to LLM with descriptive message |
Recommended for most cases |
|
All errors sent to LLM |
All errors sent to LLM |
Let LLM recover from all errors |
|
All errors propagate |
All errors propagate |
Strict mode, fail fast |
Default Handler (recommended):
val agent = OpenAIAgent
.synchronous(OpenAI.fromEnv, "gpt-4o-mini")
.maxIterations(5)
.tools(myTool)
.exceptionHandler(ExceptionHandler.default) // This is the default, can be omitted
.build
The default handler:
Propagates
IOExceptionandInterruptedException(system-level errors that typically can’t be recovered)Sends to LLM all other exceptions with descriptive error messages, allowing the agent to retry or adjust
Send All to LLM:
val agent = OpenAIAgent
.synchronous(OpenAI.fromEnv, "gpt-4o-mini")
.maxIterations(5)
.tools(myTool)
.exceptionHandler(ExceptionHandler.sendAllToLLM)
.build
All errors are converted to messages and sent to the LLM, giving it maximum opportunity to recover.
Propagate All (Strict Mode):
val agent = OpenAIAgent
.synchronous(OpenAI.fromEnv, "gpt-4o-mini")
.maxIterations(5)
.tools(myTool)
.exceptionHandler(ExceptionHandler.propagateAll)
.build
All errors immediately terminate the agent loop by propagating the exception. Use this for strict error handling where any failure should stop execution.
Custom Handler:
val customHandler = new ExceptionHandler {
def handleToolException(toolName: String, exception: Exception): Either[String, Exception] =
exception match {
case e: MyRecoverableException =>
Left(s"Recoverable error in $toolName: ${e.getMessage}")
case other =>
Right(other) // Propagate
}
def handleParseError(
toolName: String,
rawArguments: String,
parseException: Exception
): Either[String, Exception] =
Left(s"Invalid arguments for $toolName - please check the schema")
}
val agent = OpenAIAgent
.synchronous(OpenAI.fromEnv, "gpt-4o-mini")
.maxIterations(5)
.tools(myTool)
.exceptionHandler(customHandler)
.build
Return Left(message) to send the error to the LLM and continue the loop, or Right(exception) to propagate and terminate.