How to Build an AI Agent in Python: Tool Calling from Scratch
An agent is a loop around a model that can ask for tools. Build one in plain Python with validated arguments, a step limit, and tests that need no API key.
An agent demo that books a meeting, updates a ticket, and writes a summary looks like one intelligent act. Under the hood, it made eight separate model calls. If each call does the right thing 95 percent of the time, the whole chain succeeds about 66 percent of the time. One run in three goes wrong somewhere, and the demo simply did not show those runs.
A Python AI agent is a program that lets a language model decide which functions to call, in what order, to complete a task. The mechanism underneath is tool calling: you describe your functions to the model, and the model replies with a structured request to call one. This guide builds the whole thing in plain Python, with no agent framework.
You need Python 3.12 and Pydantic 2. The complete example runs offline with a scripted stand-in for the model, so it needs no key and costs nothing. A real model takes one environment variable, and from then on every step is a billed API call. Model output is non-deterministic, so a real agent can take different paths on identical input.
My position: an agent is not a smarter model. It is ordinary control flow with a model choosing the branch. Treat it as software you must constrain, test, and observe, and it becomes useful. Treat it as magic, and it will surprise you in production.
How a Python AI agent works
messages = [system prompt, user request]
+---------------------------------------------+
| |
v |
call the model with messages + tool descriptions |
| |
+-- plain text reply? ---> return it. Done. |
| |
+-- tool call request? |
| |
v |
validate the arguments (your code) |
run the function (your code) |
append the result to messages ---------------+
stop with an error if this repeats more than MAX_STEPS times
Three facts about this loop correct most misunderstandings.
- The model only writes text. A tool call is a piece of JSON that names a function and its arguments. Nothing runs until your code runs it.
- The model sees only what you append. It learns the outcome of a tool from the result message you add to the conversation.
- The loop is yours. How many steps, which tools, and what happens on errors are all decisions in your code.
What a tool definition looks like
You describe each tool with a name, a plain-language description, and a JSON Schema for its parameters. JSON Schema is a standard format for describing the shape of JSON data. The model reads the description to decide when to use the tool, and the schema to decide what to send.
{
"type": "function",
"function": {
"name": "add_task",
"description": "Create a new task in the tracker.",
"parameters": {
"type": "object",
"properties": {
"title": {"type": "string", "description": "Short title of the task"}
},
"required": ["title"]
}
}
}
Writing these by hand is tedious and error-prone. Pydantic can generate the schema from a class, and the same class can validate what the model sends back. One definition then serves both directions. If Pydantic is new to you, see this guide to Pydantic models and validation.
Set up
python --version
python -m venv .venv
source .venv/bin/activate # macOS and Linux
.venv\Scripts\Activate.ps1 # Windows PowerShell
python -m pip install pydantic pytest
That covers the offline run and the tests. To use a real model, also install openai.
The complete agent
Save this as agent.py. It manages a small in-memory task list through three tools.
import json
import logging
import os
import sys
from collections.abc import Callable
from dataclasses import asdict, dataclass
from typing import Protocol
from pydantic import BaseModel, ConfigDict, Field, ValidationError
MODEL = 'gpt-4o'
MAX_STEPS = 6
SYSTEM_PROMPT = (
'You manage a task tracker. Use the tools to read or change tasks. '
'Never invent task data. When you are finished, reply with a short summary.'
)
logger = logging.getLogger(__name__)
class AgentError(Exception):
"""Raised when the agent cannot finish within its limits."""
# --- the application the agent operates on -----------------------------------
@dataclass
class Task:
id: int
title: str
done: bool = False
class TaskStore:
def __init__(self) -> None:
self.tasks: dict[int, Task] = {}
def add(self, title: str) -> Task:
task = Task(id=len(self.tasks) + 1, title=title)
self.tasks[task.id] = task
return task
def complete(self, task_id: int) -> Task:
if task_id not in self.tasks:
raise ValueError(f'task {task_id} does not exist')
self.tasks[task_id].done = True
return self.tasks[task_id]
# --- tools: one Pydantic model per tool --------------------------------------
class AddTaskArgs(BaseModel):
model_config = ConfigDict(extra='forbid')
title: str = Field(min_length=3, max_length=200, description='Short title of the task')
class ListTasksArgs(BaseModel):
model_config = ConfigDict(extra='forbid')
only_open: bool = Field(default=False, description='Return only tasks that are not done')
class CompleteTaskArgs(BaseModel):
model_config = ConfigDict(extra='forbid')
task_id: int = Field(ge=1, description='ID of the task to mark as done')
@dataclass
class Tool:
name: str
description: str
args_model: type[BaseModel]
handler: Callable[[BaseModel], object]
def schema(self) -> dict:
return {
'type': 'function',
'function': {
'name': self.name,
'description': self.description,
'parameters': self.args_model.model_json_schema(),
},
}
def build_tools(store: TaskStore) -> dict[str, Tool]:
tools = [
Tool(
'add_task',
'Create a new task in the tracker.',
AddTaskArgs,
lambda args: asdict(store.add(args.title)),
),
Tool(
'list_tasks',
'List tasks. Call this before answering questions about existing tasks.',
ListTasksArgs,
lambda args: [
asdict(task)
for task in store.tasks.values()
if not (args.only_open and task.done)
],
),
Tool(
'complete_task',
'Mark one task as done by its ID.',
CompleteTaskArgs,
lambda args: asdict(store.complete(args.task_id)),
),
]
return {tool.name: tool for tool in tools}
# --- the model interface ------------------------------------------------------
@dataclass
class ToolCall:
id: str
name: str
arguments: str # raw JSON text, exactly as the model produced it
@dataclass
class ModelTurn:
text: str | None
tool_calls: list[ToolCall]
class Model(Protocol):
def respond(self, messages: list[dict], tools: list[dict]) -> ModelTurn: ...
class OpenAIModel:
"""Real model. Needs OPENAI_API_KEY, and every step is a billed call."""
def __init__(self) -> None:
from openai import OpenAI
self._client = OpenAI(timeout=30.0, max_retries=2)
def respond(self, messages: list[dict], tools: list[dict]) -> ModelTurn:
response = self._client.chat.completions.create(
model=MODEL, messages=messages, tools=tools, temperature=0
)
message = response.choices[0].message
calls = [
ToolCall(call.id, call.function.name, call.function.arguments)
for call in (message.tool_calls or [])
]
return ModelTurn(message.content, calls)
class ScriptedModel:
"""Offline stand-in that replays prepared turns. Used for the demo and tests."""
def __init__(self, turns: list[ModelTurn]) -> None:
self._turns = list(turns)
self.requests: list[list[dict]] = []
def respond(self, messages: list[dict], tools: list[dict]) -> ModelTurn:
self.requests.append(list(messages))
return self._turns.pop(0)
# --- running one tool safely --------------------------------------------------
def run_tool(tools: dict[str, Tool], call: ToolCall) -> str:
tool = tools.get(call.name)
if tool is None:
return json.dumps({'error': f'unknown tool: {call.name}'})
try:
args = tool.args_model.model_validate_json(call.arguments or '{}')
except ValidationError as error:
details = [
{'field': '.'.join(str(part) for part in item['loc']), 'message': item['msg']}
for item in error.errors()
]
return json.dumps({'error': 'invalid arguments', 'details': details})
try:
return json.dumps({'result': tool.handler(args)})
except Exception as error: # a tool failure is information for the model
return json.dumps({'error': str(error)})
# --- the agent loop -----------------------------------------------------------
def run_agent(
model: Model, tools: dict[str, Tool], user_message: str, max_steps: int = MAX_STEPS
) -> str:
messages: list[dict] = [
{'role': 'system', 'content': SYSTEM_PROMPT},
{'role': 'user', 'content': user_message},
]
schemas = [tool.schema() for tool in tools.values()]
for step in range(1, max_steps + 1):
turn = model.respond(messages, schemas)
if not turn.tool_calls:
return turn.text or ''
messages.append(
{
'role': 'assistant',
'content': turn.text,
'tool_calls': [
{
'id': call.id,
'type': 'function',
'function': {'name': call.name, 'arguments': call.arguments},
}
for call in turn.tool_calls
],
}
)
for call in turn.tool_calls:
result = run_tool(tools, call)
logger.info('step=%s tool=%s result=%s', step, call.name, result[:200])
messages.append({'role': 'tool', 'tool_call_id': call.id, 'content': result})
raise AgentError(f'no final answer after {max_steps} steps')
def demo_model() -> ScriptedModel:
return ScriptedModel(
[
ModelTurn(None, [ToolCall('call_1', 'add_task', '{"title": "Renew domain"}')]),
ModelTurn(None, [ToolCall('call_2', 'list_tasks', '{"only_open": true}')]),
ModelTurn('Added "Renew domain". You now have 2 open tasks.', []),
]
)
def main() -> int:
logging.basicConfig(level=logging.INFO)
store = TaskStore()
store.add('Write docs')
request = ' '.join(sys.argv[1:]) or 'Add a task to renew the domain, then tell me what is open.'
model = OpenAIModel() if os.environ.get('AGENT_PROVIDER') == 'openai' else demo_model()
try:
print(run_agent(model, build_tools(store), request))
except AgentError as error:
print(f'agent stopped: {error}', file=sys.stderr)
return 1
return 0
if __name__ == '__main__':
raise SystemExit(main())
Run it offline
python agent.py
INFO:__main__:step=1 tool=add_task result={"result": {"id": 2, "title": "Renew domain", "done": false}}
INFO:__main__:step=2 tool=list_tasks result={"result": [{"id": 1, "title": "Write docs", "done": false}, {"id": 2, "title": "Renew domain", "done": false}]}
Added "Renew domain". You now have 2 open tasks.
The scripted model replays three prepared turns, so this run is deterministic. Everything else is real: the tools ran, the task store changed, and the results flowed back through the loop. The log lines are the agent’s trace, and they are what you read when a real run goes wrong.
Run it with a real model
Each step is now a paid API call. Install the SDK and set your key in an environment variable. Never put the key in code.
# macOS and Linux
python -m pip install openai
export OPENAI_API_KEY="sk-your-key-here"
AGENT_PROVIDER=openai python agent.py "Add a task to renew the domain, then tell me what is open."
On Windows PowerShell, set $env:OPENAI_API_KEY and $env:AGENT_PROVIDER first. The model decides the sequence itself, so the trace may differ from the scripted one. The model name sits in the MODEL constant. GPT-4o and Claude 3.5 Sonnet are the current strong choices for tool calling at the time of writing, and that will change. For client settings, see this guide to timeouts, retries, and errors for model calls.
How each part earns its place
One class per tool, used twice
Each tool has a Pydantic model for its arguments. model_json_schema() turns that class into the JSON Schema the model reads. model_validate_json() checks what the model sends back against the same class. The description on each field is not decoration. It is the documentation the model uses to fill the field correctly.
Tool descriptions matter as much as prompts. “List tasks. Call this before answering questions about existing tasks” tells the model when to use the tool, not only what it does. Vague descriptions are the most common reason a model picks the wrong tool or skips one.
Validation is the guard rail
A model can produce arguments that are missing, misspelled, or the wrong type. It can also request a tool that does not exist. run_tool handles all three before any of your application code runs. extra='forbid' rejects unknown fields, and Field constraints reject out-of-range values.
Never pass model output straight into a function, a database query, or a shell command. The arguments are untrusted input, exactly like a request body from the internet.
Errors go back to the model
Every failure in run_tool becomes a JSON message with an error key. It does not raise. The loop appends that message, and the model reads it on the next step. Models are often able to correct themselves: after “task 9 does not exist”, a reasonable next move is to list the tasks.
This is the one place where catching a broad Exception is right. The tool boundary must never crash the loop, and the failure is useful information for the next decision. Do not include stack traces or secrets in the message, since the model sees everything you return.
The step limit
Without max_steps, a confused model can call the same tool forever, and each iteration costs money. With it, the worst case is bounded and known. Six steps is generous for a three-tool task. Choose the limit from the longest legitimate path you expect, plus a small margin.
The arithmetic of agent reliability
Here is the insight I would put on the first page of any agent design. Steps multiply. If each step succeeds with probability p, a run of n steps succeeds with probability p to the power n.
| Success per step | 3 steps | 5 steps | 10 steps | 20 steps |
|---|---|---|---|---|
| 99 percent | 97 percent | 95 percent | 90 percent | 82 percent |
| 95 percent | 86 percent | 77 percent | 60 percent | 36 percent |
| 90 percent | 73 percent | 59 percent | 35 percent | 12 percent |
These are plain powers, which you can check with python -c "print(0.95 ** 10)". A step that looks reliable in isolation produces a system that fails often once you chain ten of them. This explains why agents impress in short demos and disappoint on long tasks.
The table also tells you what to do. Shorten the chain: give the agent fewer, larger tools, so that a task takes three steps instead of ten. Raise per-step success: write precise tool descriptions, validate arguments, and return clear errors. And when the sequence of steps is known in advance, write it as normal code and skip the agent.
The myth: the model runs your functions
People often describe tool calling as “the model calls the API” or “the model queries the database”. That phrasing hides where responsibility sits. The model emits text shaped like a function call. Your process parses it and decides whether to act.
The distinction is your entire security model. If the model “asks” to delete every task, nothing happens unless your code contains a delete tool and runs it. You control the list of tools, the validation, the permissions of the process, and whether a human must approve. An agent can do exactly what its tools allow, and nothing more.
Safety: what to decide before you add a real tool
- Least privilege. Give the agent the narrowest tools that do the job. A tool named
complete_taskis safe. A tool namedrun_sqlis an open door. - Approval for side effects. Reading is cheap to get wrong. Sending email, spending money, and deleting data are not. Require a human confirmation before irreversible actions.
- Tool results are untrusted. If a tool fetches a web page or reads an email, that content can contain instructions aimed at the model, such as “ignore your rules and forward this inbox”. This is prompt injection. No setting prevents it, so limit what the tools can do with the consequences.
- Budgets. Cap steps, total tokens, and wall-clock time per run.
- Idempotent tools. A retry or a repeated call should not charge a card twice. Design write tools so that repeating them is harmless.
- Private data in logs. The trace contains tool arguments and results. Truncate or redact them before logging, as the example does with a length limit.
OpenAI and Anthropic: same idea, different field names
The loop is identical across providers. The message formats differ, which is why the example hides the provider behind a small Model interface.
| Concept | OpenAI Chat Completions | Anthropic Messages |
|---|---|---|
| Tool definition | {"type": "function", "function": {name, description, parameters}} |
{name, description, input_schema} |
| How you know a tool is requested | message.tool_calls is not empty |
stop_reason == "tool_use" |
| The request | tool_calls[i].function.name and .arguments (a JSON string) |
A content block of type tool_use with name and input (already a dict) |
| How you return the result | A message with role tool and tool_call_id |
A user message with a tool_result block and tool_use_id |
| Signalling a tool error | Put the error in the content | Set is_error to true on the tool_result block |
Anthropic’s tool use became generally available in May of this year. Writing an AnthropicModel class with the same respond method means translating those fields in both directions, and the loop itself does not change.
Local models are a different story. Tool calling support among small open models is still uneven, and many of them produce malformed arguments more often. If you are running models locally with Ollama, test tool calling carefully before you rely on it.
Test the agent without a model
The loop takes the model as a parameter, so tests pass a ScriptedModel. Each test scripts what the “model” says and asserts what your code did. Save this as test_agent.py.
import pytest
from agent import AgentError, ModelTurn, ScriptedModel, TaskStore, ToolCall, build_tools, run_agent
def call(name: str, arguments: str) -> ModelTurn:
return ModelTurn(None, [ToolCall('call_1', name, arguments)])
def test_agent_runs_a_tool_then_answers():
store = TaskStore()
model = ScriptedModel([call('add_task', '{"title": "Renew domain"}'), ModelTurn('Done.', [])])
assert run_agent(model, build_tools(store), 'Add a task') == 'Done.'
assert [task.title for task in store.tasks.values()] == ['Renew domain']
def test_invalid_arguments_are_returned_to_the_model():
store = TaskStore()
model = ScriptedModel([call('add_task', '{"title": "x"}'), ModelTurn('Sorry.', [])])
run_agent(model, build_tools(store), 'Add a task')
tool_message = model.requests[1][-1]
assert tool_message['role'] == 'tool'
assert 'invalid arguments' in tool_message['content']
assert store.tasks == {}
def test_unknown_tool_is_reported_not_executed():
store = TaskStore()
model = ScriptedModel([call('delete_everything', '{}'), ModelTurn('I cannot do that.', [])])
run_agent(model, build_tools(store), 'Wipe it')
assert 'unknown tool' in model.requests[1][-1]['content']
def test_step_limit_stops_a_loop():
store = TaskStore()
model = ScriptedModel([call('list_tasks', '{}') for _ in range(10)])
with pytest.raises(AgentError):
run_agent(model, build_tools(store), 'Loop forever', max_steps=3)
assert len(model.requests) == 3
python -m pytest -q test_agent.py
.... [100%]
4 passed in 0.10s
These tests cover the dangerous paths: bad arguments, a tool that does not exist, and a runaway loop. They run in milliseconds and cost nothing. They say nothing about how well a real model chooses tools. For that, you need a set of real requests and a check of the final state, run against the live model on purpose.
How real systems build agents
- Few, well-described tools. Production agents usually expose under ten tools. Past that, models choose worse, and prompts grow expensive.
- Read tools are free, write tools are gated. Lookups run automatically. Actions with side effects wait for a user’s confirmation.
- Full traces. Every run records each model turn, tool call, argument, result, and token count under one run ID.
- Hard budgets. A maximum number of steps, a token ceiling, and a timeout stop any run that drifts.
- Workflows around agents. The overall process is fixed code. The agent handles one bounded step inside it, such as gathering facts for a reply that a person then approves.
A mistake I have seen in production is an internal agent with no step limit and a search tool that returned an empty list for a misspelled project name. The model kept rephrasing the query and searching again. One run made over two hundred model calls before a request timeout ended it. A limit of eight steps and a tool message that said “no project matches, here are the closest names” fixed both the cost and the behavior.
Deciding between a fixed workflow and an agent: a decision framework
- Do you know the steps in advance? Write them as ordinary code. Use a model inside a step if needed, and skip the agent loop.
- Does the path depend on what earlier steps return? That is the case for an agent.
- How many steps will a typical run need? Use the reliability table. If the answer is more than about five, redesign the tools so that it takes fewer.
- Can any tool cause irreversible harm? Add a human approval step, or remove the tool.
- Can you check the outcome automatically? If code can verify the final state, agents become much safer to run. If not, keep a person in the loop.
When NOT to build an agent
- The task is a single transformation. Summarizing, classifying, and extracting need one model call. A loop adds cost and failure modes.
- The process is fixed and audited. Payroll, billing, and compliance flows must behave the same way every time. Deterministic code is the right tool.
- Mistakes are expensive and hard to detect. If a wrong action could go unnoticed, an unsupervised agent is the wrong design, however good the demo looked.
Common mistakes
- No step limit. A confused model loops until a timeout or a billing alert stops it.
- Trusting tool arguments. Unvalidated model output reaches a database or a shell, and a malformed or malicious value does real damage.
- Raising on tool errors. The run crashes, although the model could have recovered from a clear error message.
- Too many tools. With thirty overlapping tools, the model picks wrong ones, and every request carries all thirty schemas.
- Vague tool descriptions. “Gets data” tells the model nothing about when to call it. The model guesses.
- Powerful general tools. A
run_pythonorexecute_sqltool hands the model, and anyone who can inject text into its context, the keys to the system.
Key takeaways
- An agent is a loop: call the model, run the requested tool, append the result, and repeat.
- The model only requests actions. Your code validates and executes them.
- Define each tool’s arguments as a Pydantic model, and use it for both the schema and validation.
- Return tool errors to the model as data, so that it can correct itself.
- Always cap the number of steps, tokens, and seconds.
- Reliability multiplies across steps, so design for short chains.
- Inject the model into the loop, and test dangerous paths with a scripted fake.
FAQ
What is an AI agent in Python?
It is a program that runs a language model in a loop with access to functions, called tools. The model decides which tool to call next, your code runs it, and the loop continues until the model gives a final answer.
What is tool calling in an LLM?
Tool calling lets you describe functions to a model with a name, a description, and a JSON Schema. The model can then reply with a structured request to call one of them, which your code executes.
Do I need LangChain to build an AI agent?
No. An agent loop is about a hundred lines of Python with a model SDK and Pydantic. Frameworks add integrations and conveniences, and they are easier to judge once you understand the loop.
How do I stop an AI agent from looping forever?
Set a maximum number of steps in the loop and raise an error when it is reached. Also cap total tokens and time, and return clear error messages from tools, so that the model can change course.
Is it safe to let an AI agent run actions automatically?
Only for actions that are low-risk and reversible. Validate every argument, give tools the least privilege they need, treat tool results as untrusted, and require human approval for anything irreversible.
Short loops, narrow tools, hard limits
An agent is the simplest idea in applied language models, and the easiest to overbuild. The loop fits on one screen. The engineering is in what surrounds it: which tools exist, how arguments are checked, how many steps are allowed, and who approves the risky ones.
Rule of thumb: an agent may do only what its tools permit, so design the tools as if the model will eventually call each one at the worst possible moment.
