The Anatomy of an AI Agent: LLM, Tools, Memory, and the Control Loop

Inside the anatomy of an AI agent: the control loop, model calls, tool execution, state transitions, memory, and the failure modes each component adds.

Executive Summary: An AI agent runs as a loop where each iteration is one state transition: context in, decision out, tools run, observations return. This post breaks down the roles of the model, tools, and state, explains why the state schema matters so much, and shows how to debug a misbehaving agent by reading its trajectory instead of the prompt.

Strip away the branding and every AI agent is the same machine: a loop that connects a model to tools, state, and a stopping condition. Frameworks differ in how they package the machine. They cannot remove any of its parts. Understand the parts and you can debug any agent. Skip them and every framework becomes a black box that fails mysteriously.

This article dissects the anatomy of an AI agent: what enters the loop, what happens inside each iteration, and how decisions, actions, observations, and state interact until the run stops. It is the component-level companion to what an AI agent actually is.

What you learn here applies equally to LangGraph, CrewAI, Microsoft Agent Framework, and a hand-rolled loop, because all four implement the same anatomy. The vocabulary differs. The machine does not. Some vendors call the whole package around the model an agent harness; the AI agent harnesses explainer covers that framing, and this article opens up the loop inside it.

Observable Behavior First

Before looking inside, look at what an agent produces from the outside. A task enters. A sequence of actions happens. A result leaves. The only durable evidence of what happened in between is the trajectory: an ordered record of decisions, tool calls, and observations.

That trajectory is the primary debugging artifact in production. Every serious agent system emits it: as a trace in LangSmith or Langfuse, as OpenTelemetry spans, or as structured logs. When someone says their agent “sometimes does the wrong thing,” the first question is not which model they use. It is: show me the trajectory of a bad run.

The Control Loop

The loop is the agent. Everything else is a component the loop coordinates. Six things happen in every iteration, in this order:

            +----------------------------+
            |        AGENT LOOP          |
            |                            |
   Task --->|  1. Assemble context       |
            |  2. Model call             |
            |  3. Decision?              |
            |     |-- final answer --->  |---> Result
            |     +-- tool calls         |
            |  4. Execute tools           |
            |  5. Observations            |
            |  6. Update state            |
            |       (back to step 1)      |
            +----------------------------+
  1. Assemble context. Build what the model sees: system instructions, the task, working state, relevant memory, and the tool contracts.
  2. Model call. Send the context and the tool schemas to the model.
  3. Decision. Parse the response: a final answer ends the loop, tool calls continue it.
  4. Execute tools. Run each tool call through validation, authorization, and timeout enforcement.
  5. Observe. Convert tool results into observations the model can reason over.
  6. Update state. Advance the working state, checkpoint if needed, check budgets, and loop.

Each step is a state transition, and each transition is a place where the run can fail. Reading an agent as a table of transitions is what makes it engineerable.

Transition Trigger What changes What must be true
Assemble Loop starts an iteration Context built from state and memory Sensitive data filtered, size within budget
Decide Model responds A decision recorded: answer or tool calls Tool calls parse against their schemas
Execute Tool call validated An effect happens outside the model Authorization passed, timeout set
Observe Tool returns Observation appended to context Result validated and size-bounded
Update Observations applied Working state advances, checkpoint written State schema version respected
Stop Condition checked Loop ends with answer or escalation Stopping conditions exist and are enforced

The lineage of this shape is the ReAct pattern: reasoning traces interleaved with actions, with each action feeding evidence back into the next decision. The original paper is ReAct: Synergizing Reasoning and Acting in Language Models, and it remains the clearest description of why interleaving beats deciding everything up front.

The Model Call: Decisions Without Execution

Each iteration sends the model four things: instructions, the task and working state, the recent trajectory as context, and the contracts of the tools it may call. The model returns one of two shapes: a final answer, or one or more structured tool calls.

Two properties of the model call drive most design mistakes. First, the model can emit a tool call that does not parse: a hallucinated tool name, missing required arguments, or arguments of the wrong type. Your code must validate every call against the tool schema before execution. Second, the model sees only what you assembled. If context assembly is wrong, the best model decides badly, and no prompt fixes bad state.

Tools and Observations: Where the Loop Touches Reality

Tool execution happens in your process, with your credentials, inside your trust boundary. That is why execution is a separate transition from decision: it is where authorization, validation, and timeouts live. Tool contracts, schemas, and result handling get a full treatment in tool use and function calling, and provider mechanics are documented in guides such as the OpenAI function calling documentation.

Observations deserve their own discipline. A raw tool response is not an observation until it is validated, trimmed, and shaped for the model. A 40 kilobyte JSON dump is not an observation. It is a context-overflow event waiting for the next iteration. Production systems cap observation size, extract the fields the task needs, and convert failures into structured error observations.

That last point decides whether your agent recovers or dies. When a tool fails, the failure should usually travel back to the model as an observation: which call failed, why, and what the options are. The model can then retry with corrected arguments or choose a different route. Crashing the run on the first tool error wastes the loop’s main advantage: the ability to adapt mid-run.

State: What the Loop Remembers

Working state is the structured data the loop carries between iterations: the task, the trajectory so far, tool results, step counts, and flags. It is not the prompt. The prompt is assembled from state each iteration. The state is the source of truth that survives the iteration.

Persistence comes in layers, and conflating them causes most “my agent forgot” bugs:

Layer Lifetime Examples Owner
Working state One run Trajectory, step count, flags The loop
Conversation state One session Messages across turns Session store
Long-term memory Across sessions Preferences, learned facts Memory system
Checkpoints Crash-safe snapshots Serialized working state Checkpointer

The state schema is the contract between iterations. Keep it explicit and typed, because implicit state (facts buried inside prompt strings) is state no code can read, resume, or test. The example below is illustrative: wire the model and tools to your provider and systems.

from dataclasses import dataclass, field
from typing import Any, Callable

@dataclass
class AgentState:
    task: str
    history: list[dict[str, Any]] = field(default_factory=list)
    steps: int = 0
    done: bool = False

def run_one_iteration(state: AgentState, model,
                      tools: dict[str, Callable[..., str]],
                      max_steps: int = 10) -> AgentState:
    """Advance the loop by one iteration."""
    if state.steps >= max_steps:  # budget checked in code, not by the model
        state.done = True
        state.history.append({"role": "system", "content": "step budget exhausted"})
        return state
    response = model(state.history, tools=tools)
    state.steps += 1
    if response.tool_calls:
        for call in response.tool_calls:
            result = tools[call.name](**call.arguments)  # validate and authorize first
            state.history.append({"role": "tool", "name": call.name, "content": result})
    else:
        state.done = True
        state.history.append({"role": "assistant", "content": response.content})
    return state

What the loop persists and retrieves across sessions is a separate discipline: write policies, retrieval policies, and freshness. That is the agent memory model, and what the model sees per call is budgeted in context window management. Crash-safe snapshots and resumable runs are covered in agent state persistence.

Stopping Conditions: How the Loop Ends

A loop that relies on the model to stop is a loop that will one day not stop. Bounded execution lives in the orchestrator, and the budget check belongs at the top of every iteration, before any spend.

  • Final answer. The natural end: the model responds without tool calls.
  • Step budget. Caps iterations. Prevents infinite retry cycles.
  • Token budget. Caps total context and generation. Prevents context-growth spirals.
  • Wall-clock timeout. Caps run duration. Protects request handlers and queues.
  • Duplicate-action detection. Blocks the same tool call with the same arguments twice in a row. Prevents stuck loops that look like progress.
  • Escalation. Hands the run to a human when confidence or policy requires it.

When a budget fires, the run should fail loudly, not quietly: a structured error, a trace event, and a user-facing message that says what happened. Silent truncation produces agents that answer confidently after being cut off mid-investigation.

Failure Modes at Each Transition

Because the loop is a sequence of transitions, failures have addresses. This is what makes agents engineerable despite the nondeterministic component in the middle.

Failure Transition hit First symptom Mitigation
Hallucinated tool name or arguments Decide Tool lookup error Validate every call against its schema before execution
Tool timeout Execute Run hangs on one step Per-tool timeouts, bounded retries
Malformed or oversized result Observe Model confusion in the next iteration Validate, trim, and shape observations
Runaway loop Update Step and token counts climb Step, token, and time budgets
Context overflow Assemble Provider error mid-run Summarize, prune, compress history
Duplicate side effect Execute The same action happens twice Idempotency keys before execution
Stale memory Assemble Confident wrong answers Freshness policy, write validation

Every row is a different subsystem, which is why agent debugging is a systems discipline rather than a prompt discipline. The trajectory tells you which transition failed; the transition tells you which subsystem owns the fix.

How Real Systems Do This

Production implementations of this anatomy converge on the same habits:

  • Checkpoint at every iteration. The working state serializes after each transition, so a crash resumes from the last checkpoint instead of restarting the task and repeating side effects.
  • Cap and shape observations. Tool results are truncated, projected to needed fields, and validated before entering context. Raw API payloads never enter the loop.
  • Return tool errors as observations. The model sees what failed and why, then adapts. The run dies only when budgets or policy say so.
  • Emit a trace event per transition. Assemble, decide, execute, observe, update, stop. Debugging becomes reading a timeline instead of guessing.
  • Enforce budgets in the orchestrator. The loop checks step, token, and time budgets before each model call, not after the damage.

Decision Framework

Ask these questions about any agent you are building or reviewing, in this order:

  1. Where is the state schema? If the answer is “inside the prompt strings,” stop. Write the typed state contract first.
  2. What validates a tool call before execution? Schema validation and authorization must happen before any effect.
  3. What caps observation size? If nothing does, you have scheduled a context overflow.
  4. Which budgets run before spend? Step, token, and time limits must execute before the model call, not after it.
  5. How do tool errors reach the model? As structured observations, or as exceptions that kill the run. The first keeps the loop adaptive.
  6. Which transitions emit trace events? If fewer than all six, debugging will be archaeology.
  7. What does the user see when a budget fires? A clear message beats silent truncation.

When NOT to Use This

  • A single model call solves the task. No loop, no state machine, no tools: call the model with good context and return the result. Most LLM features are exactly this.
  • A fixed workflow solves the task. If the transitions are known, run them in code. The anatomy above buys adaptability you do not need, and pays for it in testing and observability.
  • A framework already matches your needs. When you need durable graph execution with checkpoints, use LangGraph’s primitives instead of rebuilding the loop. Hand-rolling the anatomy is for learning and for control, not for proving something.
  • You cannot yet define the state schema. The loop needs a contract between iterations. If the state is undefined, the design is not ready, and writing loop code will not fix that.

Common Mistakes

  • Implicit state in prompt strings. The price: runs that cannot be resumed, tested, or inspected, because the only copy of the truth is text nobody parses.
  • Treating tool errors as exceptions. In practice: runs that die on the first timeout instead of adapting, which discards the loop’s core advantage.
  • Unbounded observations. The result: context overflow on the third or fourth iteration, at the exact moment the task gets interesting.
  • Relying on the model to stop. The damage: the first long run that circles, and a bill that demonstrates why budgets belong in code.
  • No checkpointing. The fallout: a crash at step nine of ten restarts the whole task, repeating tool side effects the user already suffered.
  • Debugging by prompt tweaking. The cost: superstitious iteration. Reading the trajectory of transitions finds the failing step; changing words in the dark does not.

Key Takeaways

  • An agent iteration is a state transition: assemble, decide, execute, observe, update, stop.
  • The model decides. Your system executes, and the gap between those two facts is where validation, authorization, and security live.
  • Observations are engineered artifacts: validated, size-bounded, and shaped for the model. Raw tool output is not an observation.
  • Explicit, typed state is the contract between iterations. Implicit state in prompt strings is the root of unresumable, untestable agents.
  • Stopping conditions belong to the orchestrator: final answer, step budget, token budget, timeout, duplicate detection, escalation.
  • Every failure has a transition address, and the trajectory is the map. Trace all six transitions, not just the model calls.

FAQ

What are the main components of an AI agent?

Five: a model that decides, tools that act, state that survives between iterations, an orchestrator that runs the loop and enforces budgets, and stopping conditions that end the run. Memory and guardrails layer onto the same skeleton, and frameworks differ only in how they package it.

How does an AI agent control loop work?

Each iteration assembles context from state, calls the model with tool contracts, parses the decision, executes any tool calls with validation and authorization, converts results into observations, updates state, and checks stopping conditions. The loop repeats until a final answer arrives or a budget fires.

What is an observation in an agent loop?

An observation is a validated, size-bounded representation of a tool result that enters the model’s context so it can decide the next step. It is not the raw tool output. Production systems trim, project, and shape observations, because oversized or malformed results degrade the next decision.

Does the AI agent execute the tools itself?

No. The model emits tool calls as structured decisions. Your orchestration code validates them, checks authorization, and executes the tool in your process with your credentials. This separation is what makes tool permissions and audit trails possible.

What is the difference between agent state and agent memory?

State is the structured data the current run carries between iterations. Memory is what the system persists and retrieves beyond the current run: across turns or across sessions. State is the loop’s working set; memory is a retrieval problem with write and freshness policies.

Conclusion

The anatomy is stable even when the ecosystem is not. Models change every quarter, frameworks change every release, but the loop, the transition table, and the state contract are the same machine underneath all of them. Engineers who can read that machine can move between frameworks without starting over, because they debug transitions rather than syntax.

The discipline that pays off is unglamorous: typed state, validated calls, bounded observations, enforced budgets, and a trace of all six transitions. None of it demos well. All of it survives production.

When an agent fails, the answer is in the transitions. The prompt is the last place to look, not the first.

Last updated on 5 October 2026

Share this article

Leave a Reply

Your email address will not be published. Required fields are marked *