What Is an AI Agent? Definitions, Components, and How It Differs from a Chatbot
What is an AI agent? A practical definition: the components inside an agent system, how agents differ from chatbots, and when the label means something.
Ask five vendors what an AI agent is and you will get five answers, four of which describe a chat box. The word has drifted from a useful engineering concept into a marketing label. That drift becomes expensive when you need to decide whether to build one, buy one, or skip the idea entirely.
This article gives you a working definition you can defend in a design review. You will learn what an AI agent is, which components sit inside one, how it differs from a chatbot and from a plain LLM call, and where the boundaries of the term actually sit.
The short version: an AI agent is a software system built around a model that makes decisions in a loop. The model chooses actions. Your code executes them. The results return as observations, and the loop continues until a defined stopping condition fires. If a system does not do that, it is something else, and naming it correctly is the first step to designing it correctly.
What Is an AI Agent?
There is no ISO definition of an AI agent, and you should treat any claim of a single official definition with suspicion. Definitions vary by community. Researchers classify agents by degrees of autonomy. Product teams use the word for anything with a chat interface. The definition that survives engineering review is behavioral.
An AI agent is a software system in which a model decides at runtime which actions to take toward a goal. The system executes those actions, feeds the results back to the model as observations, and continues until a defined stopping condition is met.
Three questions turn that definition into a test you can apply to any product that claims to be an agent:
- Does it decide the next step at runtime? The sequence of steps is not fully hardcoded. The model chooses what happens next based on what it observes.
- Does it act on anything outside the conversation? It calls tools: APIs, databases, file systems, browsers. Text output alone does not qualify.
- Does it stop on defined conditions? A final answer, a step budget, a token budget, or a timeout. A system that only stops when the user closes the tab is not designed.
A system with all three is an agent. A system with none is a chatbot. Most confusion in the market comes from products that have one or two. A chatbot with a retrieval step still decides nothing and acts on nothing. A workflow with an LLM inside one step follows a hardcoded sequence, so the model is a component, not a decider.
| Property | Earns the agent label | Does not earn the label |
|---|---|---|
| Decisions | Model picks the next step each iteration | Transitions fixed in code, or no decisions at all |
| Actions | Executes tools with real effects | Produces text only |
| Stopping | Explicit budgets and stopping conditions | Runs until interrupted |
| State | Working state updated and persisted per run | Only a chat transcript |
Where the agent loop comes from
The loop behind this definition has a lineage. The ReAct paper (Yao et al., 2022) showed that interleaving reasoning traces with actions outperforms either alone on question answering and decision tasks, and most production agent loops are descendants of that pattern. The original paper is worth reading before you adopt any framework that implements it: ReAct: Synergizing Reasoning and Acting in Language Models.
The Components of an AI Agent
An agent is a small system with clear parts. Each part is a design decision you own, and each part gets a dedicated article in this series. The table below is the map.
| Component | What it does | The design question you own |
|---|---|---|
| Model | Decides the next action or final answer | Which model, and what happens when it is wrong |
| Tools | Turn decisions into actions on real systems | What is exposed, with what permissions and validation |
| Memory | Holds state beyond the current context window | What is persisted, retrieved, and expired |
| Orchestrator | Runs the loop, enforces budgets and stopping conditions | Who controls execution: you or the framework |
| Guardrails | Validate inputs, outputs, and tool arguments | What is blocked, corrected, or escalated to a human |
Two of those components decide most of your architecture. Tools define what the agent can affect, so they define your security boundary and your failure surface. The orchestrator defines whether runs are bounded, traceable, and resumable. The component-level deep dive, including how state and observations flow between parts, is in the anatomy of an AI agent. Tool contracts are covered in tool use and function calling.
AI Agent vs Chatbot: The Difference That Matters
A chatbot is a conversation interface over a model. You send text, it returns text. It may retrieve documents to sound informed. It cannot check an order, cancel a subscription, or open a ticket. An agent can, because it has tools, state, and a loop.
| Dimension | Chatbot | AI agent |
|---|---|---|
| Core behavior | Replies to messages | Completes tasks through a loop |
| Decisions | None, or one model call per turn | Chooses the next action each iteration |
| Actions | Text only | Calls tools with side effects |
| State | Conversation history | Run state, checkpoints, persisted memory |
| Stopping | Ends the turn, waits for the next message | Stops on conditions the system defines |
| Failure impact | A wrong or unhelpful reply | A wrong action, spend, or data change |
| Security surface | Prompt content | Tool permissions, injection, data access |
| How you evaluate | Reply quality, satisfaction | Task success, trajectory, cost per task |
The failure-impact row is the one to remember in a design review. A chatbot that hallucinates produces a bad answer. An agent that hallucinates a tool argument produces a bad action, and actions are harder to undo than sentences. That asymmetry is why agents need permissions, approval gates, and audit trails that chatbots never required.
Chatbot: message --> model --> reply
Agent: task --> decide --> tool call --> execute
^ |
+---- observation ---+
|
stopping condition --> result
Chatbots remain the right choice for many products. If the user’s goal is information or conversation, actions add cost, latency, and risk without benefit. The escalation question, when a plain conversation stops being enough, is answered in when to use an agent and when a workflow is enough.
AI Agent vs LLM vs Workflow
Three neighbors get confused with the agent, and separating them prevents most misdesigns.
- An LLM is a component. One call maps input text to output text. It decides nothing about your system; your code decides when to call it and what to do with the result.
- A workflow is a system where LLMs and tools run through predefined code paths. The pipeline owner wrote every transition in advance. Workflows are predictable, testable, and often the better choice.
- An agent is a system where the model dynamically directs its own process and tool usage to reach a goal.
That three-way distinction matches the framing Anthropic published in their engineering guidance on building effective agents, and it matches what production systems look like in practice. The workflow-versus-agent decision deserves its own treatment, because it is the decision that saves or wastes the most engineering time: read when to use an agent before committing to a loop.
What an AI Agent Is Not
Cleaning up the label is as useful as defining it. An AI agent is not any of the following:
- Not the model. The model is the decision component. The agent is the system: loop, tools, state, and guardrails. Swapping models changes quality, not architecture.
- Not a prompt. A clever system prompt does not create tools, budgets, or stopping conditions. It creates text that asks the model to behave, which is a suggestion, not a control.
- Not a framework. LangGraph, CrewAI, and Microsoft Agent Framework implement loops. They do not make your tool permissions, evaluation, or failure policy appear by magic.
- Not autonomy by default. Autonomy is a spectrum you choose: from suggest-only, to act-with-approval, to act-and-report. The default for anything irreversible should be approval first.
A Concrete Example
Consider a support agent for an online store. A customer writes: “My order is late and I want a refund.”
- The orchestrator assembles context: the message, the customer’s recent orders from a lookup tool, and the refund policy from retrieval.
- The model decides the first action: call the order status tool. It cannot invent the answer, because the status lives outside its weights.
- The tool returns: shipped 12 days ago, carrier shows no movement. That observation re-enters the context.
- The model decides the next action: check the refund policy conditions, then prepare a refund up to a policy limit.
- The refund tool is write-capable, so the system requires approval above a threshold. A human confirms. The refund executes.
- The model drafts the final reply with the refund confirmation. The stopping condition fires: final answer delivered, four steps used of a ten-step budget.
Every step in that example is testable. The tool calls are logged, the approval is an audit event, and the outcome is measurable: refund issued, customer informed, cost known. That testability is what separates an engineered agent from a demo. The full run mechanics, including how observations update state, are covered in the anatomy of an AI agent.
Failure Modes of Poorly Defined Agents
Vague definitions produce concrete failures. The recurring ones:
- Label inflation. A pipeline gets called an agent, so the team adds a model where deterministic code belonged. The result costs more and fails more often.
- Missing stopping conditions. The loop continues while the model keeps finding “one more thing” to check. The result is runaway token spend.
- Over-permissioned tools. Because “the agent needs access,” it gets broad credentials. The result is a large blast radius for every hallucination.
- Unmeasurable success. Nobody defined what a correct run looks like, so every regression is discovered by users instead of tests.
How Real Systems Do This
The label becomes clearer when you inspect real products instead of landing pages.
- Coding assistants are the clearest agents in wide use. They read a repository, decide which files to inspect, propose edits, run tests, and iterate on failures. Each iteration is a decision made at runtime, with tool effects and a stopping condition.
- Support automation usually pairs a chatbot front end with an agent behind it. Simple questions stay in conversation mode. The moment the user wants an action (a refund, a plan change, a ticket), the loop with tools takes over. The interface is one product; the architectures are two.
- Research assistants run a loop over search and browsing tools, then synthesize with citations. The browsing steps are the agent part; the synthesis step is a model call inside the run.
- Products marketed as agents are often workflows underneath. That is not a scandal. Fixed pipelines are easier to operate and audit. The lesson is to check behavior, not the label on the pricing page.
Decision Framework
Apply these questions to any system you are building, buying, or reviewing. Answer in order, and stop at the first no.
- Does the system choose its next step at runtime? If no, it is a workflow or a single model call. Design it as such, with the predictability that brings.
- Does it act on external systems through tools? If no, it is a chatbot or an LLM application. It needs conversation design and retrieval, not an agent architecture.
- Are stopping conditions explicit in code? If no, it is a prototype. Budgets and timeouts are what make a loop deployable.
- Are irreversible actions gated by approvals? If no, the design is unsafe regardless of what you call it.
- Can you define and measure success? If no, fix that before building anything, agent or otherwise.
Five yes answers mean you have an agent and the responsibilities that come with it: tracing, evaluation, cost control, and a permission model. That full journey starts with the complete guide to AI agent architecture and production deployment.
When NOT to Use This
- When the user only needs answers. Information requests, explanations, and drafting are conversation or single-call work. Adding tools multiplies cost and risk for no benefit.
- When the process is fixed and stable. Known pipelines with known exceptions are workflows. An agent loop adds variance to a process whose value is consistency.
- When actions are irreversible and trust is not established. Payments, deletions, and external communications need deterministic paths plus human approval. An agent can prepare the action; it should not be the only thing between a guess and a consequence.
- When success cannot be defined. “Make the user happy” is not a specification. Without an objective signal, an agent is untestable and unimprovable.
Common Mistakes
- Calling everything an agent. The consequence is architectural: teams add loops and models to fixed processes and inherit nondeterminism they did not need.
- Treating the model as the system. The result is missing infrastructure: no budgets, no state, no audit trail, and a prototype that cannot ship.
- Assuming a chat wrapper with tools is production-ready. What follows is the first long run that loops until a provider limit stops it.
- Defining agents by autonomy instead of by behavior. The outcome is over-scoping: teams build “fully autonomous” systems when a suggest-and-approve loop would have solved the problem.
- Ignoring the security shift. Actions mean permissions. The consequence of skipping that shift is an agent whose credentials define your blast radius.
- Skipping the definition conversation. The price is a team that argues about vocabulary in the design review instead of deciding state, boundaries, and metrics.
Key Takeaways
- An AI agent is a system: a model making runtime decisions inside a loop that executes tools and stops on defined conditions.
- Three properties earn the label: runtime decisions, actions with effects, and explicit stopping conditions.
- A chatbot produces text. An agent produces effects. The difference defines your failure surface and your security work.
- An LLM is a component, a workflow is a predefined path, and an agent decides its own path. All three are valid; only one needs agent engineering.
- The model is replaceable. The loop, tools, state, and guardrails are the product you maintain.
- Verify claims by behavior, not by label. Many “agents” are workflows, and many workflows should stay workflows.
FAQ
Is ChatGPT an AI agent?
In its basic chat form, no: it is a chatbot interface to an LLM, producing text without acting on external systems. When the same product browses the web, runs code, or completes multi-step tasks on your behalf, those modes behave like agents: runtime decisions, tool actions, and stopping conditions. The product contains both architectures.
Do AI agents need to be autonomous?
No. Autonomy is a design choice on a spectrum: suggest-only, act-with-approval, or act-and-report. The defining properties are runtime decisions, tool actions, and stopping conditions, not independence. Many production agents are deliberately narrow and heavily supervised, because supervision is what makes them safe to operate.
Is an AI workflow the same as an AI agent?
No. A workflow runs LLMs and tools through transitions defined in code in advance. An agent chooses its own transitions at runtime. Workflows are more predictable and cheaper to operate. Agents are more flexible and more expensive to control. Many production systems use both.
Do all AI agents use LLMs?
In current product practice, most do, because language models are the most capable general-purpose decision component available. The definition does not require them: classic software agents in robotics and games decided and acted without LLMs. This series focuses on LLM-based agents, where the new engineering challenges concentrate.
What is the difference between an AI agent and an AI assistant?
An assistant is a product surface: a voice or chat interface that helps a user. An agent is a system behavior: a loop that decides and acts. Early voice assistants were mostly not agents, because their skills were fixed paths. Modern assistants increasingly behave agent-like when they execute multi-step tasks with tools.
Conclusion
Definitions are not academic when they decide architecture. Calling something an agent commits you to a loop, tools, state, budgets, permissions, and an evaluation story. Calling it a chatbot or a workflow commits you to something cheaper and more predictable. Both commitments are respectable. The expensive mistake is paying agent costs for chatbot needs, or shipping chatbot reliability under an agent name.
Use the behavioral test before every build decision: does it decide at runtime, does it act through tools, does it stop on conditions you wrote? With those three answers, the rest of this series gives you the depth for each part: the anatomy of the loop, the tool contracts, the memory model, and the production machinery.
If you cannot say what a system decides, what it touches, and when it stops, you do not have an agent yet. You have a name.
Last updated on 3 October 2026
