Tool Use and Function Calling: Giving Agents the Ability to Act
How tool use and function calling work in AI agents: tool contracts, JSON schemas, validation, authorization, and error handling that survives production.
A model with no tools is a very fast writer. A model with tools can check an order, query a database, send a message, or break one. Tool use and function calling are where agent engineering gets real: decisions become actions, and actions have consequences.
The mechanics look simple. Describe a function. The model emits a structured call. Your code executes it. The engineering lives in everything around that exchange: the contract, the validation, the authorization, the result shape, and the failure behavior.
This article builds one coherent tool layer in Python, then stress-tests it. What happens when arguments are wrong? When a tool times out? When the result is malformed? And when the caller is not allowed? A tool layer that survives those four questions is production-ready.
The Problem: Models Generate Text, Systems Need Actions
A language model maps text to text. An application needs effects: a record read, a ticket created, a payment refunded. Function calling is the bridge. The model receives the contracts of available functions, decides which to call with which arguments, and emits a structured call instead of, or alongside, prose. Your system executes it and returns the result as an observation.
Terminology varies by vendor. OpenAI’s documentation calls it function calling. Anthropic and most frameworks say tool use. The mechanism is identical: structured intent from the model, execution in your code. This article uses “tool” for the capability and “tool call” for the structured decision.
One distinction prevents a category of design errors: exposing a tool is a definition decision. Running it is an execution decision. The model can only request. Your registry decides. Everything in this article exists to make that gap deliberate.
Tool Contracts: The Schema Is the Interface
A tool contract has four parts, and each has a different consumer:
| Contract part | Who reads it | What goes wrong if it is bad |
|---|---|---|
| Name | Your registry and the model | Collisions, silent misrouting |
| Description | The model | Wrong tool chosen, wrong situations, missed preconditions |
| Parameter schema | Your validator and the model | Hallucinated arguments, crashes, corrupt writes |
| Execution metadata | Your gateway | Over-authorization, unsafe retries, missing audit |
The description is the part teams underinvest in, and it is the part the model reads. State what the tool does, when to use it, when not to use it, and whether it has side effects. Anthropic’s agent engineering guidance makes the same point with a rule worth keeping: write tool descriptions like a great docstring for a new colleague, and test how the model uses them instead of assuming. Their full guidance, including tool ergonomics, is at building effective agents.
Execution metadata is the part no vendor tutorial shows: which role may call the tool, whether the tool is idempotent, and what its blast radius is. It does not go to the model. It goes to your gateway, which is the subject of tool sandboxing and permission models.
How Function Calling Works: The Full Lifecycle
Every provider wraps the same six-step lifecycle. Learn it once and read any SDK fluently.
- Register. Build tool contracts and expose them to the model alongside the task context.
- Decide. The model returns zero or more structured calls: a name, an arguments object, and a call identifier.
- Parse. Decode the arguments JSON. Malformed JSON becomes an error observation, not an exception.
- Validate. Check the arguments against the contract’s JSON Schema. Reject unknown keys, wrong types, and missing required fields.
- Authorize and execute. Check the caller’s role, then run the handler with a timeout. Your code, your credentials, your boundary.
- Observe. Shape the result, append it to context, and let the next iteration decide.
Model decision: call lookup_order(order_id="ORD-12345")
|
v
Parse JSON arguments
|
v
Validate against JSON Schema ---- failure ---> error observation
|
v
Authorization check --------- denied ------> error observation + audit log
|
v
Execute handler with timeout -- timeout -----> error observation, maybe retry
|
v
Shape result, cap size, return as observation
Steps 3 through 5 are where demo code and production code diverge. Demos assume the model emits perfect calls. Production assumes it sometimes does not, and treats every failure as recoverable input. The loop that consumes these observations is covered in the anatomy of an AI agent.
A Complete Tool Layer in Python
The following implementation is deliberately framework-free. It shows the moving parts you would otherwise find hidden inside an SDK. The model client interface is illustrative; wire it to your provider. For a complete, runnable version against a real provider, see how to build an AI agent in Python with tool calling, and for schema-validated model output in general, see structured output from LLMs with Pydantic.
import json
from dataclasses import dataclass
from typing import Any, Callable
@dataclass(frozen=True)
class ToolSpec:
name: str
description: str
parameters: dict[str, Any] # JSON Schema
handler: Callable[[dict[str, Any]], Any]
allowed_roles: tuple[str, ...] = ("agent",)
idempotent: bool = False
class ToolRegistry:
"""Holds contracts, validates calls, and executes handlers."""
def __init__(self) -> None:
self._tools: dict[str, ToolSpec] = {}
def register(self, spec: ToolSpec) -> None:
if spec.name in self._tools:
raise ValueError(f"duplicate tool: {spec.name}")
self._tools[spec.name] = spec
def contracts_for_model(self) -> list[dict[str, Any]]:
return [
{
"name": spec.name,
"description": spec.description,
"parameters": spec.parameters,
}
for spec in self._tools.values()
]
Validation is a separate function, not a decorator or a convention, so it is testable on its own. This checker is intentionally minimal: production systems should use a full JSON Schema validator, which also enforces enums, formats, and nested structure. Like every snippet that follows, it builds on the previous block and reuses its imports, so read them as one module.
TYPE_MAP = {
"string": str,
"integer": int,
"number": (int, float),
"boolean": bool,
}
def validate_arguments(schema: dict[str, Any],
args: dict[str, Any]) -> tuple[bool, str]:
"""Minimal schema check. Use a full validator in production."""
for key in schema.get("required", []):
if key not in args:
return False, f"missing required argument: {key}"
for key, rule in schema.get("properties", {}).items():
if key in args:
expected = TYPE_MAP.get(rule.get("type", ""))
if expected and not isinstance(args[key], expected):
return False, f"wrong type for argument: {key}"
return True, "ok"
Execution belongs to the registry, and every failure path returns a structured observation instead of raising. A tool layer that throws on bad input kills runs; a tool layer that reports bad input keeps the loop alive.
def execute(self, name: str, arguments: dict[str, Any],
actor_role: str) -> dict[str, Any]:
spec = self._tools.get(name)
if spec is None:
return {"error": f"unknown tool: {name}"}
if actor_role not in spec.allowed_roles:
return {"error": "caller not authorized for this tool"}
ok, reason = validate_arguments(spec.parameters, arguments)
if not ok:
return {"error": f"invalid arguments: {reason}"}
try:
result = spec.handler(arguments)
except Exception as exc: # a tool must never crash the loop
return {"error": f"tool failed: {exc}"}
return {"result": result}
Registering a tool makes the contract style concrete. Note the description: what it does, when to use it, and the fact that it is read-only. That last clause changes model behavior measurably.
def lookup_order(args: dict[str, Any]) -> dict[str, Any]:
"""Illustrative body. Replace with a real data query."""
return {"order_id": args["order_id"], "status": "shipped", "eta_days": 2}
REGISTRY = ToolRegistry()
REGISTRY.register(ToolSpec(
name="lookup_order",
description=(
"Fetch the current status of one customer order by its id. "
"Use when the user asks where an order is. Read-only, no side effects."
),
parameters={
"type": "object",
"properties": {
"order_id": {
"type": "string",
"description": "Order identifier, such as ORD-12345.",
}
},
"required": ["order_id"],
},
handler=lookup_order,
idempotent=True,
))
Wiring the registry into a loop is short, because the loop’s job is only to shuttle decisions and observations. The step budget lives here, in code, not in the model’s judgment. The model client signature is illustrative.
def handle_tool_calls(model, task: str, actor_role: str = "agent",
max_steps: int = 5) -> str:
"""Decide, execute, observe, repeat. Illustrative model client."""
response = model(task, tools=REGISTRY.contracts_for_model())
for _ in range(max_steps):
if not response.tool_calls:
return response.content
observations = []
for call in response.tool_calls:
try:
arguments = json.loads(call.arguments)
except json.JSONDecodeError:
outcome = {"error": "malformed JSON in tool call"}
else:
outcome = REGISTRY.execute(call.name, arguments, actor_role)
observations.append({"tool": call.name, **outcome})
response = model(task, context=observations,
tools=REGISTRY.contracts_for_model())
return "stopped: step budget exhausted"
Tool Results: Design What Comes Back
The result is not for the database. It is for the model’s next decision. Three rules keep observations useful:
- Project, do not dump. Return the fields the task needs, not the full API payload. A 40 kilobyte response in context is a future context overflow, and it buries the signal the model needs.
- Cap the size. Enforce a maximum observation length and truncate with a marker. The model can ask for more with a follow-up call if needed.
- Fail in structure, not prose. Error observations should be structured: what was called, what failed, why, and what the model can do about it. “Tool failed” teaches nothing; “order_id not found, check the format with the search tool” teaches recovery.
Authorization Happens Before Execution
Exposing a tool to the model is an act of trust, and that trust should be configured, not implied. The model cannot enforce permissions on itself. A system prompt that says “only refund eligible orders” is a suggestion to the component most likely to hallucinate. The registry is the enforcement point.
Three mechanisms cover most needs:
- Role tiers on every tool. Read-only tools get the default role. Write tools require an elevated role, and irreversible tools require an approval workflow on top. The human-in-the-loop pattern covers the approval design.
- Idempotency flags that govern retries. Only idempotent tools may be retried automatically. Every write tool should accept an idempotency key so a retry after a timeout cannot double the effect.
- Audit events on every execute. Who called, what tool, what arguments, what result code. When something goes wrong at 2 a.m., the audit trail is the difference between an incident report and a mystery.
Least privilege and capability isolation, including what a sandbox actually restricts, get the full treatment in tool sandboxing and permission models.
Failure Handling for Tool Calls
Every failure in the table below becomes an observation. None of them become exceptions. That single rule is what keeps a loop alive through bad hours.
| Failure | Where it is detected | Response |
|---|---|---|
| Unknown tool name | Registry lookup | Error observation; fix the description or register the tool |
| Malformed JSON arguments | Parse step | Error observation asking for corrected arguments |
| Schema violation | Validation | Error observation naming the violated field |
| Authorization denial | Role check | Error observation plus audit event; never hint at a bypass |
| Timeout | Execution wrapper | Bounded retry if idempotent; otherwise escalate |
| Handler exception | Execution wrapper | Error observation; open the circuit breaker on repeats |
Two cautions complete the policy. Never retry a non-idempotent tool automatically; a timeout means you do not know whether the effect happened. And when one tool fails consistently, stop calling it: circuit breaking exists because a model that “knows” a tool should work will keep trying it.
Testing Tools Without the Model
The model is the only untestable part of a tool layer. Everything else should have tests, because everything else is deterministic code.
def test_rejects_missing_argument() -> None:
schema = {"type": "object", "required": ["order_id"], "properties": {}}
ok, reason = validate_arguments(schema, {})
assert not ok and "order_id" in reason
def test_execute_denies_wrong_role() -> None:
outcome = REGISTRY.execute("lookup_order", {"order_id": "ORD-1"}, "guest")
assert "not authorized" in outcome["error"]
Four test classes cover the layer: validation tests with hand-written argument dicts, authorization tests with each role, execution tests with handlers that raise on purpose, and loop tests with a fake model that emits canned tool calls. Teams that skip these tests end up testing in production, with the model as the only, and most expensive, test runner.
How Real Systems Do This
- My flagship systems ship read-only first too. My engineering intelligence agent delivers its value through read-only telemetry ingestion and analysis tools with leadership-facing outputs, which is why its worst failure is a weak insight, not a broken pipeline. The write-capable tier arrives only where the task demands it.
- Tool gateways. Enterprise deployments put one service between agents and internal systems. Authorization, rate limiting, schema validation, and audit logging happen in one choke point instead of inside every tool.
- Read-only first. New agents ship with lookup tools only. Write capabilities arrive behind role upgrades and approval gates after the run history justifies them.
- Descriptions are versioned artifacts. Teams iterate tool descriptions with the same seriousness as prompts, because tool-selection errors are usually description errors. Changes ship with evaluation runs, not vibes.
- Idempotency keys on every write. Refunds, sends, and creates all accept a key so retries after timeouts stay safe.
- MCP for reuse. When several clients need the same capability, teams expose it once as an MCP server instead of writing an integration per client. The protocol and its boundaries are covered in MCP (Model Context Protocol).
Decision Framework
Before exposing any tool, answer these in order:
- What decision does this tool enable? If the model cannot meaningfully choose to call it, call it yourself in code and skip the model entirely.
- Is it read-only or a write? Writes need a role tier, an idempotency key, and usually an approval gate.
- Could the description be misread? Run scenarios and watch tool selection. Fix the contract before blaming the model.
- Which result fields does the task need? Project to those. Everything else is context pollution.
- What happens on timeout and failure? Define the observation and the retry policy before the first incident.
- Who may call it, and where is that logged? If there is no audit answer, there is no audit.
- Will more than one client need it? Then standardize the contract instead of writing it three times.
When NOT to Use This
- The task needs no actions. Classification, extraction, and drafting are single structured-output calls. Wrapping them in a tool layer adds failure modes without adding capability.
- The arguments are already known. If your code knows the order id, call the API directly. A model in the middle of a deterministic call adds cost, latency, and a new way to be wrong.
- The tool has unbounded blast radius and no audit trail. Expose it only after the permission model exists. Capability without logging is an incident on layaway.
- The call sits on a latency-critical path. Hot paths belong to code. Model-mediated tool calls are for decisions, not for throughput.
Common Mistakes
- One-line descriptions. “Gets order.” The result: the model calls the order tool for refund questions, shipping questions, and everything containing the word order.
- Raw payload dumps. The damage: context overflow and degraded decisions on the fourth tool call, blamed on the model.
- No authorization layer. The fallout: one over-permissioned tool turns prompt confusion into a security event.
- Exceptions that crash the loop. The cost: runs that die on the first flaky dependency, with no recovery and no trace of what the model learned.
- Retrying non-idempotent tools. What you get: duplicate refunds and duplicate emails, discovered by the recipients.
- Tool sprawl. Forty overlapping tools with fuzzy boundaries. Where it lands: selection errors that no model upgrade fixes, because the contracts are the problem.
Key Takeaways
- A tool is a contract plus an execution path. The model requests; your code decides and executes. That gap is your security model.
- Descriptions are written for the model, schemas for your validator, and execution metadata for your gateway. Three audiences, one contract.
- Validate against the JSON Schema before execution, and authorize before validation. Both produce structured error observations, never exceptions.
- Design results for the next decision: project fields, cap size, fail in structure.
- Retry only idempotent tools. Give every write an idempotency key, and circuit-break tools that fail repeatedly.
- Test the layer without the model. Everything except the model is deterministic and should have tests.
FAQ
What is function calling in AI agents?
Function calling is the mechanism that lets a model emit a structured request to run a named function with typed arguments, instead of only producing text. Your system receives the request, validates it, executes the function, and returns the result to the model as an observation it can reason over.
What is the difference between tool use and function calling?
They describe the same mechanism with different emphasis. Function calling is the vendor term for the structured request. Tool use is the broader engineering term that also covers how you design, authorize, execute, and observe the capability. If a blog post treats them as rivals, it is arguing about vocabulary, not architecture.
How do you validate AI agent tool calls?
Validate the decoded arguments against the tool’s JSON Schema before execution: check required fields, types, and allowed values. Return failures as structured error observations so the model can correct its arguments. Never pass unvalidated model output directly to a handler.
Can an AI agent call any API?
No. An agent can only call tools you exposed, with the permissions you configured. The model emits requests; your registry or gateway enforces authorization. Treat any system where the model can reach credentials directly as a design defect, not a feature.
What makes a good tool description?
What the tool does, when to use it, when not to use it, its side effects, and the argument formats with examples. Write it like a docstring for a capable new colleague, then watch real runs and fix the description where selection goes wrong.
Conclusion
Tool use is the difference between a system that talks about work and a system that does it. The mechanism is a structured call, but the engineering is a contract discipline: descriptions the model can use, schemas your validator enforces, permissions your gateway owns, and observations your loop can survive.
The investment is asymmetric in your favor. A tool layer built this way outlives model upgrades, survives provider switches, and turns most runtime failures into recoverable observations instead of crashed runs.
An agent is exactly as capable, and exactly as dangerous, as the tools you expose. The contract is where you decide which of those you are building.
Last updated on 6 October 2026
