AI Agents: The Complete Guide to Architecture, Frameworks, and Production Deployment
A complete guide to building reliable AI agents: architecture, tools, memory, frameworks, evaluation, observability, security, and production deployment.
Most teams do not fail at AI agents because the model is weak. They fail because the system around the model is missing. An agent that answers a question in a notebook is not the same product as an agent that runs in production, calls real tools, spends real money, and fails in real ways.
An AI agent is not simply an LLM with a prompt. It is a software system that reasons over a task, chooses actions, calls tools, observes results, updates state, and continues until it reaches a defined stopping condition. Every word in that sentence is an engineering commitment. Reasoning is a model call. Actions are tool executions. Observations are data you must validate. State is something you must persist and recover.
This guide is the entry point to a complete series on building AI agents. It maps the territory end to end: components, architecture, frameworks, multi-agent patterns, memory, retrieval, evaluation, observability, deployment, cost, and security. Each section links to a deeper article where you can go narrow.
Read it top to bottom if you are new to agents, or jump to the section that matches your current decision.
What Is an AI Agent?
Marketing copy implies an agent is any LLM product with a chat box. That definition will not survive contact with production. An agent is a system in which a model participates in a loop of decisions and actions. It selects a tool, your system executes the tool, the result re-enters the model context, and the loop continues until a stopping condition fires.
An AI agent is a software system that reasons over a task, chooses actions, calls tools, observes results, updates state, and continues until it reaches a defined stopping condition.
No single universal definition exists, and you should distrust anyone who claims otherwise. Vendors stretch the word to cover chat wrappers. Research papers define agents by degrees of autonomy. For engineering work, the useful test is behavioral. Does the system choose its own next step? Does that choice trigger an action outside the model, such as an API call, a database write, or a file change?
That behavioral test separates three things teams often confuse. An LLM is a component: one call, one completion. A chatbot is a product: a model behind a conversation interface with no actions. An agent is a system: a model plus tools, state, and a loop. The table below makes the boundaries concrete.
| Dimension | LLM | Chatbot | AI agent |
|---|---|---|---|
| Interaction | Single call and completion | Turn-based replies | Goal-directed run |
| Decisions | None | Fixed flow or one model call | Model picks the next step each iteration |
| Actions | Text only | Text only | Executes tools with real side effects |
| State | Request payload | Conversation history | Working state plus durable checkpoints |
| Failure impact | A bad completion | A bad reply | A wrong API call, spend, or data change |
| How you evaluate it | Answer quality | Answer quality | Task success, trajectory, cost |
Agents vs workflows
The other boundary that matters is the one between agents and workflows. A workflow runs steps in a sequence you defined in advance. An agent decides its own sequence. Many systems labeled as agents are workflows with an LLM inside one step, and that is often the right design. Anthropic drew the same line in their engineering guidance on building effective agents: workflows orchestrate models through predefined code paths, while agents dynamically direct their own processes and tool usage.
For the full definitional treatment, read what an AI agent actually is. For the component-level view, see the anatomy of an AI agent. The build-or-not decision starts with when to use an agent and when a workflow is enough.
Core AI Agent Architecture
Every agent reduces to the same skeleton, regardless of vendor or framework. Learn the skeleton first. Every framework then becomes a variation you can evaluate instead of adopt on faith.
The loop has six steps:
- Receive the task and assemble context: system instructions, conversation state, relevant memory, retrieved documents.
- Call the model with the task, the state, and the tool contracts available to it.
- Parse the decision: a final answer, a tool call, or several tool calls.
- Execute each tool call with validation, authorization, and timeouts.
- Append results as observations, update state, and check stopping conditions.
- Repeat until the agent returns a final answer or hits a budget limit.
User Request
|
v
Orchestrator (control loop)
|
+--> Model call (decides the next action)
|
+--> Tool gateway ---> External APIs, databases, SaaS
|
+--> Memory store (read and write state)
|
+--> Retrieval layer (RAG)
|
+--> Guardrails and human approvals
|
v
Stopping condition met ---> Final result
The following example is illustrative. It shows the loop with no framework, so you can see what frameworks actually manage for you. Replace the model client with your provider SDK.
import json
from typing import Any, Callable
Tool = Callable[[dict[str, Any]], str]
TOOLS: dict[str, Tool] = {}
def tool(name: str, func: Tool) -> None:
TOOLS[name] = func
def get_order(args: dict[str, Any]) -> str:
"""Illustrative lookup. Replace with a real data source."""
return json.dumps({"order_id": args["order_id"], "status": "shipped"})
tool("get_order", get_order)
def run_agent(model, task: str, max_steps: int = 10) -> str:
"""Minimal agent loop: decide, act, observe, stop."""
messages: list[dict[str, Any]] = [{"role": "user", "content": task}]
for _ in range(max_steps):
response = model(messages, tools=list(TOOLS))
if not response.tool_calls:
return response.content # stopping condition: final answer
for call in response.tool_calls:
result = TOOLS[call.name](call.arguments)
messages.append({"role": "tool", "name": call.name, "content": result})
raise RuntimeError("step budget exhausted without a final answer")
The model is a decision component, not the agent
The model chooses the next action. It does not execute anything, remember anything, or enforce anything. Treating the model as the whole agent is the root cause behind most fragile designs. The system around the model owns execution, state, and failure handling.
Tools turn decisions into actions
A tool is an executable capability exposed to the model through a contract, usually JSON Schema. The contract, the authorization to execute, and the validation of results are all yours to design. The tool use and function calling article covers the full lifecycle.
Memory and context are different problems
The context window is the model’s working set for this run. Memory is everything the system persists beyond it. Teams that treat conversation history as memory end up with agents that cannot retrieve the right state when it matters. The distinction gets a full treatment in agent memory and context window management.
Planning is an orchestration pattern, not magic
ReAct-style loops, plan-and-execute, and decomposition are orchestration patterns you control, not properties the model acquires. The planning and reasoning patterns article explains when each pattern earns its complexity.
Stopping conditions are a safety mechanism
Every production loop needs explicit stopping conditions: a final answer, a step budget, a token budget, or a timeout. An agent without a stopping condition is a loop that waits for your bill to stop it. This connects directly to reliability engineering for agents.
Human oversight belongs in the loop, not after it
Approvals, escalation, and review are control boundaries. They cost less when designed in from the start than when bolted on after the first bad tool call. See human-in-the-loop controls for agents.
Agents vs Workflows: Choosing the Right Shape
A workflow is a sequence of steps where you defined the transitions in advance. An agent is a loop where the model decides the transitions. The distinction sounds academic until the first production failure. Workflows fail predictably and are easy to debug. Agents fail creatively, in ways that include executing the wrong tool with valid-looking arguments.
| Property | Deterministic workflow | Agent loop |
|---|---|---|
| Control flow | Fixed transitions, known in advance | Model chooses the next step each iteration |
| Predictability | High, same input gives the same path | Low to medium, path varies by run |
| Cost profile | Known cost per execution | Variable, depends on steps taken |
| Failure modes | Step errors, timeouts | Bad tool choice, loops, hallucinated arguments, context overflow |
| Debugging | Logs and step status | Full trajectory tracing required |
| Best fit | Known, stable processes | Open-ended tasks with unpredictable steps |
A workflow is often the better architecture when every transition is known in advance. The judgment that matters: start with a workflow, then replace only the steps that genuinely need model judgment. Many production systems are hybrids. A deterministic pipeline routes and validates, and an agent loop handles the one open-ended step, such as investigating an incident or drafting a response from retrieved evidence.
For the full decision framework, including the escalation path from workflow to agent, read when to use an agent.
Agent Frameworks and Orchestration
You can write the loop yourself in a hundred lines, as the Python tool-calling agent built from scratch shows. Frameworks earn their place when you need durable state, checkpointing, streaming, human-in-the-loop primitives, and retry semantics that you would otherwise rebuild badly. What no framework provides: your tool authorization model, your business logic, your evaluation suite, and your security boundaries.
LangGraph: graph orchestration with durable state
LangGraph is a low-level orchestration framework and runtime for long-running, stateful agents. You define a state schema, nodes that transform state, and edges that can form cycles. Its core strengths are durable execution, checkpointing, human-in-the-loop interrupts, and the ability to mix deterministic hand-coded steps with model-driven steps in one graph. If you need fine control over the loop, this is the layer to learn. The official documentation lives at docs.langchain.com, and the dedicated guide here is LangGraph for agents: stateful, cyclic workflows.
CrewAI and Microsoft Agent Framework: higher-level abstractions
CrewAI models work as role-based crews: agents with roles, goals, and tools, plus flows for deterministic orchestration, with guardrails, memory, and structured outputs built in. On the Microsoft side, AutoGen is now in maintenance mode: it receives no new features and is community managed. Its successor is Microsoft Agent Framework, built by the AutoGen and Semantic Kernel teams, which adds graph-based workflows with checkpointing and human-in-the-loop support, built-in OpenTelemetry tracing, and MCP and A2A interoperability. These abstractions trade control for speed of assembly. The comparison, including when each fits and when none of them fit, lives in CrewAI vs AutoGen vs LangGraph: choosing a framework.
MCP: one protocol instead of N times M integrations
The Model Context Protocol (MCP) is an open standard for connecting AI applications to external systems: data sources, tools, and prompts. Without it, every client and every tool server needs a bespoke integration. With it, a tool written once works across clients that support the protocol, including major assistants and coding tools. MCP standardizes discovery and invocation, not trust. Authorization, sandboxing, and audit remain your system-level responsibilities. At the time of writing, the current specification is dated 2026-07-28. The official site is modelcontextprotocol.io, the agent-side view is MCP for AI agents, and the full protocol walkthrough is MCP explained: the Model Context Protocol guide.
| Option | Abstraction level | Best fit |
|---|---|---|
| Plain loop | You write the loop | Learning the mechanics, simple two-step agents, maximum control |
| LangGraph | Graph orchestration runtime | Stateful, cyclic workflows with checkpoints and human-in-the-loop |
| CrewAI | Role-based crews and flows | Team-style delegation with structured outputs |
| Microsoft Agent Framework | Agents plus graph-based workflows | Multi-agent systems on the Microsoft stack, and the migration path from AutoGen |
| MCP | Tool interoperability protocol | Standardizing tool access across clients and models |
Choose the abstraction level that matches the control you need over the loop. A demo does not justify a framework. A production system with resumable state and approval gates probably does not justify a hand-rolled loop either.
Multi-Agent Systems: Delegation With a Price
Multi-agent architectures split a task across specialized agents: a supervisor that routes work, a swarm of peers that hand off directly, or a hierarchy with layers of managers and workers. The patterns are real and useful. The costs are equally real: coordination overhead, duplicated context, higher token usage, longer latency, harder debugging, and failure that propagates across agents.
Adding a second agent increases coordination cost, so delegation should solve a real bottleneck. Multi-agent earns its complexity when subtasks need genuinely different prompts, tools, or permissions. It does not earn it because the diagram looks impressive. The pattern catalog is in multi-agent architectures: supervisor, swarm, and hierarchical, and the cost-benefit test is in when multi-agent beats single-agent.
Memory, Retrieval, and Context
Three different problems hide under the word “memory.” The context window holds what the model sees now. Short-term memory is the state of the current run. Long-term memory is what the system persists across sessions, retrieved selectively rather than replayed wholesale. Retrieval-augmented generation (RAG) is retrieval infrastructure: query, search, rank, assemble, ground. It feeds the loop with evidence. It is not the same thing as memory.
Memory is useful only when the system can retrieve the right state at the right time. A store that returns stale or irrelevant history degrades the model instead of helping it. The design work sits in three policies: what deserves persisting, what deserves loading, and what deserves expiring. Read agent memory: short-term, long-term, and episodic for the memory model, RAG for agents for task-adaptive retrieval, and context window management for budgeting, summarization, compression, and pruning.
Evaluation and Observability
Demo success is not production reliability. A demo shows the happy path on a handful of chosen prompts. Production requires knowing the failure rate across the full task distribution, the cost distribution, and the recovery behavior when a tool fails mid-run.
Evaluate more than final text quality. The metrics that matter for agents: task success rate, tool-selection correctness, tool argument correctness, trajectory quality, policy compliance, latency, token usage, and cost per successful task. Cost per model call is the wrong denominator. A cheap model that needs four retries to complete a task can cost more per successful task than a strong model that completes it in one pass.
An agent run is a distributed workflow, so trace it like one: a root run, agent steps, model calls, tool calls with latency, retrievals, retries, state transitions, and the final outcome. Tracing platforms such as LangSmith and Langfuse, plus vendor-neutral OpenTelemetry, make this practical. One warning before you ship anything: prompts, tool arguments, and tool results can contain secrets and personal data. Do not export all agent content to telemetry systems by default.
The measurement playbook lives in how to evaluate AI agents. The tracing stack is covered in tracing agent runs: LangSmith, Langfuse, and OpenTelemetry.
Production Deployment Architecture
Deployment is where an agent stops being a script and becomes a distributed system. The execution boundary grows from a function call into a set of services: an API gateway, an orchestrator, a model provider, a tool gateway, state and retrieval stores, and a queue for long runs.
Client
|
v
API / Gateway
|
v
Agent Orchestrator
|
+--> Model Provider
|
+--> Tool Gateway
| +--> Internal API
| +--> Database
| +--> SaaS API
|
+--> Memory / State Store
|
+--> Retrieval Layer
|
+--> Queue / Worker
|
+--> Observability
Decide early whether a run is synchronous or asynchronous. A quick question can complete inside a request-response cycle. A task that calls six tools and takes three minutes belongs in a background job with checkpointed state, a status endpoint, and idempotency on resume. Retrying a non-idempotent run after a timeout can double a side effect, such as a second refund or a second email.
The production checklist does not fit in a pillar guide, so it is split across four dedicated articles:
- Deploying AI agents: architecture and patterns
- Reliability for agents: retries, timeouts, fallbacks, and circuit breakers
- Cost control: token budgets, caching, and model routing
- Agent state persistence: checkpointing and resumable workflows
Security and Permission Boundaries
Agents extend your security perimeter. The model reads untrusted content: user input, retrieved documents, web pages, tool output. It can then act on trusted systems through your tools. That combination does not exist in a normal chatbot, and it is why agent security is an architecture problem rather than a prompt problem.
Untrusted User Input
|
v
Agent / Model
|
v
Tool Selection
|
v
Authorization Boundary
|
v
Tool Execution
|
v
External System
Prompt injection is the headline risk. Instructions smuggled into data the agent reads, such as a retrieved document or a web page, can steer it toward actions the user never requested. The OWASP Top 10 for LLM applications lists prompt injection as LLM01 and excessive agency as LLM06, which matches what practitioners see: most incidents come from over-permissioned tools, not from clever prompts alone. The reference is on the OWASP GenAI Security Project site.
Defense in depth is the only honest position: least privilege, trusted and untrusted data separation, structured tool contracts, output validation, authorization, approval gates, sandboxing, network controls, and audit logs. A system prompt is not a security boundary, because prompt content is data, not a control. An agent with unrestricted tools is not autonomous. It is underconstrained.
The attack paths and countermeasures are covered in prompt injection in agents: indirect attacks through tools and data, and the permission model is covered in tool sandboxing and permission models: least privilege for agents. Approval boundaries that put a human before irreversible actions are designed in human-in-the-loop: approvals, escalation, and guardrails.
How Real Systems Do This
My own flagship systems run these same shapes. My engineering intelligence agent ingests software delivery telemetry, analyzes execution patterns, and produces leadership-ready insights through dashboards: a read-heavy analysis loop with a hard quality bar. The product solution architect agent I architected and delivered researches target financial institutions and synthesizes business intelligence from public and enterprise sources: research and synthesis over retrieved material, which is a loop, not a chat box. Neither is a chat wrapper. Both are tools, budgets, and audit trails around a model.
Production agent systems converge on a small set of patterns, regardless of vendor.
- Support and operations agents start with read-only tools, a step budget, and human approval for any write action. The first production milestone is a defensible audit trail, not autonomy.
- Coding agents execute in sandboxed environments with network egress controls, produce diffs rather than direct commits, and pass test gates before review. The sandbox is a permission boundary, not a performance feature.
- Research and analysis agents separate browsing, which ingests untrusted content, from synthesis, which produces the answer. Citations and provenance are required, because an unsourced claim is unauditable.
- Enterprise deployments put a tool gateway between the agent and internal systems. Authorization, rate limiting, and audit logging happen in one place instead of inside every tool.
None of these patterns depend on exotic model capability. They depend on ordinary distributed-systems discipline applied to a component that happens to be nondeterministic.
Decision Framework
Work through these questions in order. Each answer narrows the architecture.
- Is the step sequence predictable? If yes, use a deterministic workflow and stop reading agent tutorials.
- Does the task need actions, not just answers? If no, a single model call with retrieval may be enough.
- Can you define success objectively? If no, you cannot evaluate the agent. Fix that first.
- What cost per successful task is acceptable? Set the budget before you pick a model or a framework.
- Which actions are irreversible? Those need approval boundaries and audit logs.
- How will you debug a failed run? If you have no tracing answer, build observability before autonomy.
- Do subtasks need different prompts, tools, or permissions? Only a yes justifies multi-agent.
When NOT to Use This
Agents are the wrong tool in at least four situations:
- The process is fully known and stable. Invoice parsing with a fixed schema is a pipeline. An agent loop adds cost, latency, and failure modes without adding capability.
- There is no objective success signal. If you cannot write a test for the outcome, you cannot iterate on the agent or detect regressions. You will ship vibes.
- Actions are irreversible and high-stakes. Payments, legal filings, and infrastructure changes need deterministic paths with human sign-off. An agent can prepare the action. It should not silently execute it.
- Latency budgets are tight. A loop with several model calls and tool round trips cannot beat a single cached call. Interactive typing-ahead UIs are a poor fit.
Common Mistakes
- Shipping the demo loop. No step budget, no cost cap, no timeout. The consequence is a run that loops until it hits provider limits, and a bill that matches.
- Treating the framework as the architecture. Frameworks manage the loop, not your tool permissions, evaluation, or failure policy. The result is a polished demo of someone else’s control flow.
- Exposing over-permissive tools. One admin-capability tool with the agent’s credentials turns a prompt problem into a security incident. What follows is blast radius.
- Confusing conversation history with memory. Replaying full history until the context window fills degrades accuracy and inflates cost. The outcome is an agent that gets worse as conversations get longer.
- Evaluating only final answers. A correct answer reached through wrong tool calls will fail at scale. The price is silent regressions that user complaints discover before your metrics do.
- Adding agents for show. Multi-agent without a delegation bottleneck multiplies tokens, latency, and debugging surface. You get a system that costs more and fails in more places.
Key Takeaways
- An agent is a control loop: reason, act, observe, update state, stop. The model is one component inside that loop, not the product.
- The engineering effort lives outside the model call: tool contracts, state, failure handling, cost limits, and permission boundaries.
- Prefer a workflow when transitions are known. Add an agent loop only where the path must be decided at runtime.
- Frameworks reduce loop boilerplate. They do not remove authorization, evaluation, or security from your scope.
- Measure cost per successful task, task success rate, and trajectory quality. Demo quality is not a metric.
- Every loop needs explicit stopping conditions: final answer, step budget, token budget, timeout.
- Least privilege is the primary security control. A system prompt is not a boundary; an authorization check is.
FAQ
What is the difference between an AI agent and a chatbot?
A chatbot produces text inside a conversation. An agent executes actions inside a loop: it calls tools, observes results, updates state, and continues until a stopping condition fires. The chatbot’s worst failure is a bad reply. The agent’s worst failure is a wrong action with side effects.
Do I need a framework like LangGraph to build an agent?
No. A minimal agent loop is a hundred lines of code. A framework earns its place when you need durable state, checkpointing, human-in-the-loop interrupts, and streaming, or when you want those primitives maintained by someone else. Start without one until you feel the specific pain it solves.
When should I use multi-agent instead of a single agent?
Only when subtasks need genuinely different prompts, tools, or permissions, or when context isolation between subtasks improves reliability. Multi-agent adds coordination cost, latency, token usage, and debugging surface. Delegation should solve a real bottleneck, not an organizational diagram.
What does an AI agent cost to run in production?
It depends on model choice, steps per task, and retry behavior, so model the metric that matters: cost per successful task, not cost per model call. Include tool API costs, infrastructure, observability, and human review time. A strong model that finishes in one pass often beats a cheap model that loops four times.
How do I stop an agent from looping forever?
Enforce explicit stopping conditions in the orchestrator: a step budget, a token budget, a wall-clock timeout, and duplicate-action detection. Make the loop raise or escalate when a budget is exhausted. Never rely on the model to stop itself.
Conclusion
AI agents are not a new kind of intelligence. They are a new shape of software: a nondeterministic decision component inside a loop that touches real systems. That shape makes the standard distributed-systems disciplines matter more, not less: state, retries, idempotency, observability, cost control, and least privilege.
The teams that succeed treat the model as a replaceable component and the loop as the product. They start with the smallest agent that solves the task, measure cost per successful task, and buy autonomy only when the workload pays for it.
The rule of thumb for the whole series: the hard part of an agent is rarely the model call. It is controlling what happens before and after the call.
Last updated on 7 October 2026
