LangGraph for Agents: Stateful, Cyclic Workflows

How LangGraph models agents as stateful, cyclic graphs: state schemas, nodes, checkpointing, human review interrupts, and when to choose it over a loop.

Executive Summary: LangGraph models an agent run as a graph of nodes over an explicit state schema, with cycles and a checkpointer that saves state at every step. This post explains how that gives you crash recovery, resumable runs, and human approval mid-graph, what LangGraph leaves to you, and when a plain loop or a higher-level framework fits better.

A hand-rolled agent loop is a hundred lines. It works until the process dies at step seven, until you need a human to approve step eight before it runs, or until the task outgrows a single request. Then you discover you were not writing a loop. You were writing a workflow engine, badly.

LangGraph is a low-level orchestration framework and runtime for building, managing, and deploying long-running, stateful agents. It models an agent as a graph: a state schema, nodes that transform state, and edges that can cycle. The model calls are just nodes inside your structure.

This article explains how LangGraph models stateful, cyclic agent workflows, what it manages for you, what it deliberately does not, and when to choose it over a plain loop or a higher-level framework. The companion comparison of LangGraph against CrewAI and AutoGen (now in maintenance mode and succeeded by Microsoft Agent Framework) lives in CrewAI vs AutoGen vs LangGraph.

The Problem LangGraph Solves

The agent control loop has properties that plain code handles poorly:

  • It is cyclic. The next step depends on the model’s decision, so control flow is a loop with branches, not a straight line.
  • It is long-running. A run with ten tool calls and an approval gate can outlive a request, a container, or a workday.
  • It must pause. Human approvals, budget checks, and escalations all need the run to stop and resume with full state intact.
  • It must be observable. A run that cannot show its path cannot be debugged.

Naive implementations fail at each property: the loop lives in one process, the state lives in memory, the pause is a hack, and the path is inferred from logs. LangGraph exists to make those four properties ordinary instead of heroic. Its own documentation is direct about the positioning: a low-level orchestration framework for mixing deterministic, hand-coded steps with LLM-driven agentic steps in one graph, with durable execution, streaming, and human-in-the-loop as core capabilities, documented at docs.langchain.com.

The Core Abstraction: State, Nodes, Edges

LangGraph models a run as a graph over three components, and the documentation compresses the mental model into one sentence worth memorizing: nodes do the work, edges tell what to do next.

  • State is a shared, typed schema representing the current snapshot of the run: typically a TypedDict, where each field can declare a reducer that defines how updates merge. A messages field with an append reducer accumulates the trajectory instead of overwriting it.
  • Nodes are functions that receive the state and return updates to it. A node can be a model call, a tool executor, a validator, or plain code. The framework does not care which.
  • Edges determine the next node: fixed transitions or conditional functions that branch on state. Cycles are allowed, which is what makes an agent loop expressible as a graph at all.

Execution proceeds in discrete super-steps, a design inspired by Google’s Pregel system. Nodes that run in parallel execute within the same super-step; sequential nodes occupy separate ones. State updates apply between super-steps, which is also where checkpointing hooks in.

          +--------+
START -->| model  |  decides: answer or tool calls
          +---+----+
              |
     conditional edge (cycle allowed)
              |
       +------+------+
       |             |
       v             v
  +--------+      +--------+
  | tools  |---->| model  |  observations re-enter state
  +--------+      +---+----+
                      |
                      v
                    END      final answer

A Minimal Graph

The smallest useful graph mirrors the official quickstart: one state schema, one node, straight edges, compiled and invoked.

from langgraph.graph import StateGraph, MessagesState, START, END

def mock_llm(state: MessagesState):
    return {"messages": [{"role": "ai", "content": "hello world"}]}

builder = StateGraph(MessagesState)
builder.add_node(mock_llm)
builder.add_edge(START, "mock_llm")
builder.add_edge("mock_llm", END)
graph = builder.compile()

result = graph.invoke({"messages": [{"role": "user", "content": "hi!"}]})

Install with pip install -U langgraph. Note what is absent: no provider lock-in, no prompt framework, no agent abstraction. LangGraph can be used without LangChain, a point its documentation makes explicitly, because its subject is orchestration, not components. If you want the component side first, LangChain 101: chains, prompts, and the first LLM app covers it.

Adding the Cycle: The Agent Loop as a Graph

An agent needs the graph equivalent of the control loop: a model node, a tools node, and a conditional edge that keeps the cycle going until the model returns a final answer. The mock stands in for your provider call, exactly as the quickstart uses one; the recursion limit acts as the step budget.

from typing import Annotated, Literal, TypedDict
from langgraph.graph import StateGraph, START, END

class AgentState(TypedDict):
    messages: Annotated[list, lambda existing, update: existing + update]

class Reply:
    """Stub provider answer: swap in your provider's message type."""
    content = "Order ORD-12345 shipped two days ago."
    tool_calls = None

def execute_pending_calls(messages: list) -> list:
    """Stub: validate and run the last message's tool calls through
    your tool layer, returning one tool message per call."""
    return []

def call_model(state: AgentState) -> dict:
    """Node: one model decision, through the stub above."""
    return {"messages": [Reply()]}

def run_tools(state: AgentState) -> dict:
    """Node: execute validated tool calls, results as tool messages."""
    return {"messages": execute_pending_calls(state["messages"])}

def route(state: AgentState) -> Literal["tools", END]:
    last = state["messages"][-1]
    return "tools" if getattr(last, "tool_calls", None) else END

builder = StateGraph(AgentState)
builder.add_node("model", call_model)
builder.add_node("tools", run_tools)
builder.add_edge(START, "model")
builder.add_conditional_edges("model", route)
builder.add_edge("tools", "model")
agent = builder.compile()

result = agent.invoke(
    {"messages": [{"role": "user", "content": "Where is order ORD-12345?"}]},
    {"recursion_limit": 25},
)

Every piece of the anatomy maps: the model node is the decision transition, the tools node is execution, the append reducer is the observation log, and recursion_limit is the stopping condition enforced by the runtime rather than by the model. When the limit is hit, the runtime raises GraphRecursionError, and the docs recommend handling budgets proactively inside the graph with a remaining-steps check that routes to a graceful end, rather than catching the exception after the fact.

Checkpointing and Durable Execution

Checkpointing is the capability that separates a graph runtime from a loop in a script. After each super-step, the state snapshot persists through a checkpointer, keyed to a thread. A run that crashes at step seven resumes from the last checkpoint instead of restarting and repeating tool side effects. A run that pauses for an approval resumes with the decision injected. The docs also cover time travel, replaying a run’s state history, which turns “what did the agent see at step four” from an archaeology project into a lookup.

The general concepts (checkpoint contents, replay safety, versioning, and side-effect discipline) are covered framework-free in agent state persistence. LangGraph is one concrete implementation of that article’s model.

Interrupts: Human-in-the-Loop as a Graph Property

LangGraph supports interrupting a graph at a node, persisting, and resuming later with a human’s decision. That is what makes approval gates from human-in-the-loop design implementable without hacks: the interrupt is not a blocked thread holding memory hostage, it is a persisted state waiting for input. The same mechanism serves escalation and budget checks.

What LangGraph Does Not Manage

The documentation is unusually candid about this, and the list matters more than the feature list:

  • Prompts and architecture. LangGraph does not abstract them. You own every prompt, every node’s behavior, and the overall design.
  • Tool authorization. The tools node executes what you give it. Roles, validation, and sandboxing remain yours.
  • Evaluation. A graph that cannot be evaluated is still a liability, framework or not.
  • Business logic. Nodes are functions; the domain rules inside them are your code and your tests.

Strengths and Limitations

Dimension Assessment
State model Explicit, typed, versionable; reducers make accumulation rules visible
Cycles and branching First-class; conditional edges replace if-else forests
Durability Checkpointing, resume, time travel as runtime features
Human-in-the-loop Interrupt and resume as graph properties, not hacks
Control granularity Maximum; you place every deterministic and agentic step
Abstraction level Low: more structure to write than role-based frameworks
Learning curve Real: state schemas, reducers, and super-steps need practice
Version sensitivity APIs evolve; documentation itself moved domains recently

Production Considerations

  • Budgets first. Set recursion_limit on every invoke, and prefer proactive in-graph budget handling to exception handling after the failure.
  • Observability. Graph structure makes traces legible, and LangSmith integrates directly for tracing and evaluation of LangGraph runs, documented at docs.smith.langchain.com; the stack is covered in tracing agent runs.
  • State schema versioning. Checkpoints serialize your state shape. When the schema changes, old threads need a migration story; the docs dedicate sections to graph migrations and backward compatibility for exactly this reason.
  • Node discipline. One job per node. Giant nodes recreate the monolith the graph was supposed to replace.
  • Testing. Nodes are functions of state, which makes them unit-testable without the model, a property the production documentation leans on.

How Real Systems Do This

  • The canonical production shape is small. A model node, a tools node, one conditional edge, budgets on. Most graphs in production are this plus a validation node and an approval interrupt, not sprawling diagrams.
  • Deterministic and agentic steps share one graph. Validation, policy checks, and enrichment run as plain-code nodes, and only the open-ended steps call the model. Mixing the two is the framework’s central design point, not an advanced technique.
  • Write tools sit behind interrupts. The run persists, a human decides, the graph resumes. Approval becomes a graph property instead of a blocked thread.
  • Checkpointers are enabled from day one. Teams that run without checkpointing in development find out about resume bugs in production, where they are most expensive.
  • Debugging reads the graph. Because paths are explicit, trajectory review maps one-to-one onto nodes and edges, which is why the graph abstraction pays for itself during incidents.

Decision Framework

  1. Must runs survive crashes or pause for humans? If yes, you need a durable runtime, and LangGraph is the leading open-source option in Python.
  2. Is your control flow branching and cyclic enough that if-else forests are unreadable? If yes, the graph formalism pays. If no, it is ceremony.
  3. Do you want loop control or delegation structure? Loop control is LangGraph. Role-based delegation and multi-agent conversation are what CrewAI and Microsoft Agent Framework optimize for; the trade-offs are in choosing a framework.
  4. Is this a two-step prototype? Write the plain loop. Adopting a runtime before you feel its absence teaches you nothing except its syntax.
  5. Can the team own state schema versioning? Checkpoints serialize your schema. If nobody owns migrations, durability becomes a liability.
  6. Does the step-level control justify the learning curve? State schemas, reducers, and super-steps take practice. Budget for it honestly.

When NOT to Use This

  • Single-call tasks. A graph around one model call is a dependency with no job.
  • Fixed, linear pipelines. Deterministic workflows belong in ordinary code or a workflow engine. The graph formalism adds concepts without adding capability.
  • When you want roles and crews, not steps. If your mental model is “a researcher agent hands work to a writer agent,” a higher-level framework fits that model directly.
  • When the team cannot maintain the state schema. Graphs with unversioned state and checkpointers enabled are deferred production incidents.

Common Mistakes

  • Monolithic nodes. One node doing retrieval, parsing, validation, and three model calls. What you get: a graph diagram that lies, and tests that cannot isolate anything.
  • No recursion limit. Where it lands: the first cyclic run becomes a cost incident, with GraphRecursionError playing the role the budget should have played.
  • Unversioned state schemas. The consequence: old threads become unreadable after the first migration, and resume quietly corrupts.
  • Business logic inside edges. Routing functions should decide, not do. What follows: invisible side effects that fire on every path evaluation.
  • Checkpointing deferred. The price: a system that looks durable and is not, discovered during the first crash.
  • Expecting the framework to be the architecture. In practice: prompts, tool contracts, permissions, and evaluation arrive late, and the graph merely organizes their absence.

Key Takeaways

  • LangGraph models an agent as a graph: typed state, nodes that transform it, edges that route it, cycles allowed.
  • Durable execution is the headline capability: checkpointing, resume, and time travel are runtime features, not weekend projects.
  • Interrupts make human-in-the-loop a graph property: persisted state waiting for a decision, not a blocked thread.
  • The agent loop is two nodes and a conditional edge. Complexity beyond that should be earned, not assumed.
  • LangGraph does not manage prompts, authorization, evaluation, or business logic. Those remain yours under any framework.
  • Set recursion limits on every invoke, and version your state schema before enabling long-lived threads.

FAQ

What is LangGraph used for?

Building stateful, cyclic agent workflows that need durable execution: checkpointed state, resumable runs, human-in-the-loop interrupts, and streaming. It is a low-level orchestration runtime, so it suits teams that want control over every step rather than a prepackaged agent abstraction.

Is LangGraph the same as LangChain?

No. LangChain is a framework of components and prebuilt agent abstractions. LangGraph is the orchestration runtime underneath: graphs, state, checkpointing, interrupts. The documentation states plainly that LangGraph can be used without LangChain, and that it focuses on agent orchestration rather than components.

What is a super-step in LangGraph?

A discrete execution step, inspired by Google’s Pregel system. Nodes that run in parallel execute within one super-step; sequential nodes occupy separate ones. State updates apply between super-steps, which is also where checkpoints are taken.

How does LangGraph checkpointing work?

A checkpointer persists the state snapshot after each super-step, keyed to a thread. Crashed runs resume from the last checkpoint, paused runs resume with injected input, and the state history supports time travel for inspecting what the agent saw at any step.

When should I not use LangGraph?

For single-call tasks, fixed linear pipelines, or when you want role-based delegation abstractions rather than step-level control. If nothing needs to survive a crash or pause for a human, a plain loop is the simpler and more honest architecture.

Conclusion

LangGraph earns its place by making the hard parts of agent loops ordinary: cycles are first-class, state is typed and versionable, runs survive crashes, and humans can pause them without hacks. The cost is a real learning curve and the discipline of owning an explicit state schema.

The framework also enforces a useful honesty. Because it abstracts so little, every design decision stays visible in the graph, and every operational responsibility stays in your scope. Teams that want that visibility adopt it comfortably. Teams that want delegation metaphors should look at the alternatives before committing.

Use LangGraph when the run must survive crashes, pauses, and scrutiny. Use plain code when it must merely run.

Last updated on 3 October 2026

Share this article

Leave a Reply

Your email address will not be published. Required fields are marked *