Planning and Reasoning: ReAct, Plan-and-Execute, and Tree-of-Thought

AI agent planning and reasoning patterns: how ReAct, plan-and-execute, and tree-of-thought work, and when each one earns its complexity in production.

Executive Summary: Reasoning picks an agent’s next step and planning picks its route. This post compares ReAct, plan-and-execute, and Tree of Thoughts with the cost and failure profile of each, explains when reflection helps, and recommends choosing the cheapest pattern your task can tolerate and proving it with trajectory evaluation.

Ask an agent to “think harder” and you have changed nothing. What changes behavior is the orchestration pattern around the model: how the loop decides, when it plans, how it recovers, and where it verifies.

Planning and reasoning patterns are engineering structures, not model moods. ReAct interleaves reasoning with action. Plan-and-execute separates the route from the driving. Tree of Thoughts searches branches and backtracks. Reflection critiques its own output and tries again.

This article explains each pattern, what it costs, and when it earns its complexity. The organizing rule throughout: a pattern is an observable orchestration decision. If you cannot see the pattern in the trace, it does not exist.

None of these patterns require a specific model feature or hidden reasoning exposure. They are loop shapes you control, evaluated by the trajectories they produce, which is why the measurement machinery from how to evaluate AI agents applies to every section below.

Planning vs Reasoning: Two Different Questions

The two words get used interchangeably, and the confusion is architectural, not just grammatical.

  • Reasoning answers: what is the next step, given what I know right now? It operates inside a loop iteration. Its input is the current context; its output is one decision.
  • Planning answers: what sequence of steps should this run follow? It operates above the iterations. Its input is the task; its output is a route, which the loop then walks and, when reality disagrees, revises.

You do not need a planning model to have reasoning, and you do not need visible reasoning to have a plan. A fixed workflow is a plan with no runtime reasoning. A ReAct loop is runtime reasoning with no upfront plan. The interesting systems are explicit about which question they are answering where, because the agent control loop hosts both.

Why Patterns Matter More Than Prompts

A prompt can ask the model to plan, reflect, or explore. If nothing in the system enforces, stores, or branches on that request, the behavior exists only until the model gets bored of it. A pattern, by contrast, has structure the trace can show: a plan object with steps, a critique pass with a stored reflection, a branch queue with evaluated candidates.

That test keeps the vocabulary honest. “We use Tree of Thoughts” means there is search code that maintains branches. “We use reflection” means there is a critique step whose output changes the next attempt. Anything else is prompt decoration, and prompt decoration cannot be tested, budgeted, or debugged. It also cannot be evaluated, which matters because pattern choice is ultimately a cost-quality trade you must measure.

ReAct: Reason and Act, One Step at a Time

ReAct is the default shape of production agents. Each iteration produces a small piece of reasoning and one action, the system executes the action, the observation re-enters the context, and the loop continues until a final answer or a budget fires.

        +-----------+
Task -->|  Reason   |  what do I know, what is next?
        +-----+-----+
              |
              v
        +-----------+
        |    Act    |  one tool call
        +-----+-----+
              |
              v
        +-----------+
        |  Observe  |  result re-enters context
        +-----+-----+
              |
              +-- repeat until final answer or budget

The original paper (Yao et al., 2022) showed why interleaving beats deciding everything up front: reasoning traces help the model track and update its plan, while actions pull fresh evidence from external sources. On question answering and fact verification, the interleaved approach reduced the hallucination and error propagation of reasoning-only baselines by grounding each step in retrieved evidence. The paper is short and worth reading directly: ReAct: Synergizing Reasoning and Acting in Language Models.

Use ReAct when the route depends on observations. Its weaknesses are equally structural: latency grows with steps, the loop can wander when no plan exists, and without budgets it can circle. The wandering is not a model defect. It is what a plan-shaped problem looks like inside a step-shaped pattern.

Plan-and-Execute: Decide the Route First

Plan-and-execute splits the work in two. A planner produces an explicit, inspectable list of steps. An executor walks the steps, often with a small ReAct loop inside each one. When a step fails or the world disagrees with the plan, a replanner revises the route.

from dataclasses import dataclass

@dataclass
class StepResult:
    ok: bool
    output: str
    error: str = ""

def plan_and_execute(model, task: str, max_steps: int = 12) -> str:
    """Illustrative skeleton: plan, execute, replan on failure."""
    steps: list[str] = model.make_plan(task)
    done: list[str] = []
    for _ in range(max_steps):
        if not steps:
            return model.final_answer(task, done)
        step = steps[0]
        result = execute_step(model, step)
        if result.ok:
            done.append(f"{step}: {result.output}")
            steps = steps[1:]
        else:
            steps = model.replan(task, done, step, result.error)
    return "stopped: step budget exhausted"

def execute_step(model, step: str) -> StepResult:
    """One step may itself be a small tool loop. Illustrative."""
    try:
        output = model.run_step(step)
    except Exception as exc:
        return StepResult(ok=False, output="", error=str(exc))
    return StepResult(ok=True, output=output)

The model client interface is illustrative; wire it to your provider and tools. What is not optional is the structural point: the plan is an object. It can be shown to a human for approval before execution, persisted, resumed, and revised, which makes plan-and-execute the natural pattern for long-horizon work.

What you buy: foresight on long tasks, parallel execution of independent steps, and a reviewable artifact before anything irreversible happens. What you pay: a planner call up front, replan complexity when reality diverges, and the risk of stale plans being walked faithfully into failure. A plan is a hypothesis about the world. The replanner is what makes it a correctable hypothesis.

Tree of Thoughts: Search Over Branches

Tree of Thoughts (ToT) generalizes linear reasoning into search. Instead of committing to one next thought, the system generates several candidates, evaluates them, explores the promising branches, and backtracks from dead ends.

The evidence for the pattern is strong in a narrow domain. In the paper’s Game of 24 experiments, GPT-4 with chain-of-thought prompting solved 4 percent of instances, while the tree-of-thought search solved 74 percent (Yao et al., 2023). The difference is the problem shape: combinatorial search with a verifiable success condition, where lookahead and backtracking are the whole game: Tree of Thoughts: Deliberate Problem Solving with Large Language Models.

Production use is narrower than the enthusiasm suggests. Full ToT multiplies model calls per decision and needs evaluation and pruning logic you must write and debug. What most production systems actually run is the cheap approximation: generate N candidates, score them with an objective verifier or a judge, keep the best. That is best-of-N selection, it inherits much of ToT’s benefit, and it fits inside a normal step budget. Reserve full search for problems where success is verifiable in code and the branching structure is real.

Supporting Patterns: Decomposition, Routing, Reflection, Verification

Four smaller patterns appear inside and between the big three. They are cheap, and they do most of the useful work in production systems.

Decomposition

Split a task into subtasks before solving them. It is the engine inside plan-and-execute and the justification for multi-agent delegation: subtasks that need different prompts, tools, or permissions are natural split points. The delegation mechanics are covered in multi-agent architectures.

Routing

Classify the input, then dispatch to a specialized handler. It is the cheapest pattern here, often one fast model call or even an embedding similarity check, and it is frequently the only “intelligence” a workload needs. Many systems that think they need an agent need a router plus three fixed paths, the distinction drawn in when to use an agent.

Reflection

Generate, critique, retry. Reflexion (Shinn et al., 2023) made the case concretely: agents that verbally reflect on task feedback and keep those reflections in an episodic buffer improved materially on subsequent attempts, including 91 percent pass@1 on HumanEval against an 80 percent GPT-4 baseline. The paper is at Reflexion: Language Agents with Verbal Reinforcement Learning.

The caveat is the feedback source. Reflection works when the critique is grounded in something external: test results, compiler errors, a verifier, user corrections. A model critiquing its own output with no external signal tends toward self-congratulation, and self-congratulation does not survive evaluation.

Verification

An independent check on the output: unit tests for code, schema validation for structured data, a judge for tone, a second pass with a checklist. Verification is the only pattern in this article that adds objectivity instead of adding effort. When a checker exists for your task, use it before reaching for any reasoning pattern: it is cheaper than search and more honest than reflection.

Pattern Comparison

Pattern What it decides Extra cost Earns its complexity when
ReAct The next action, each iteration Baseline The route depends on observations
Plan-and-execute The route, up front and on replan Planner plus replanner calls Long multi-step tasks, parallel steps, reviewable plans
Tree of Thoughts Which branch to pursue Multiplied calls plus search logic Combinatorial problems with code-verifiable success
Routing Which handler One classifier call Inputs divide into known categories
Reflection Whether to retry, and how A critique pass per attempt Quality matters and real feedback exists
Verification Accept or reject One checker call An objective check exists at all

The table’s quiet rule: verification and routing cost one call. ReAct costs a loop. Plan-and-execute costs a loop plus a planner. ToT costs a loop per branch. Ascending the table is a cost decision, and the only honest way to make it is with evaluation results on your own tasks.

How Real Systems Do This

  • Most production agents are bounded ReAct. Step budgets, token budgets, duplicate detection, and a small tool set. The pattern is ordinary; the budgets are the product.
  • Long-horizon jobs use plan-and-execute with visible plans. Research reports, migrations, and batch investigations persist the plan, show it for approval where stakes are high, and revise it on failure instead of walking it blindly.
  • Support systems lead with routing. One classifier call splits traffic, fixed paths handle the known majority, and a loop handles the rest. The router often removes 90 percent of the load before any agent is involved.
  • Coding agents lean on verification, not self-belief. Compilers, linters, and test suites are the objective critics. The strongest reasoning pattern in coding agents is act, verify with a real checker, fix, repeat.
  • Quality-critical generation runs best-of-N with a judge. Generate candidates, verify against criteria, keep the best, log the scores. Full tree search is rare; the approximation captures most of the value.

Decision Framework

Work down this list and stop at the first pattern that fits:

  1. Does the route depend on what tools discover? If yes, start with bounded ReAct. It is the baseline that everything else extends.
  2. Is the task long, multi-step, or parallelizable? Plan first. Show the plan where approval matters, and make failure trigger a replan rather than a shrug.
  3. Does an objective checker exist? Use verification before any reasoning pattern. It is the cheapest quality lever in this entire article.
  4. Is the search space combinatorial with cheap verification? Use best-of-N selection. Escalate to full tree-of-thought search only when the approximation measurably fails.
  5. Do inputs split into known categories? Route. Do not loop over what a classifier can dispatch.
  6. What are the latency and cost ceilings? Multiply each pattern’s model calls by your volume. The right pattern is the one that fits the arithmetic, not the diagram.

When NOT to Use This

  • A single call solves the task. Patterns exist to manage multi-step uncertainty. One-step tasks need good prompts and good outputs, not search.
  • There is no objective success signal. Without one, reflection flatters itself and search maximizes the wrong thing. Define success before buying complexity.
  • Interactive latency is the product. Multi-pass patterns do not fit sub-second interfaces. Cache, route, and keep the loop off the hot path.
  • The workflow already encodes the plan. Re-planning in the model what code already guarantees adds variance on top of determinism, which is a strict loss.

Common Mistakes

  • Cargo-culting complexity. Tree search on linear tasks. The consequence is multiplied cost with no measurable gain, discovered in the invoice review.
  • Plans without replanning. The result is faithful execution of a stale hypothesis, which is worse than no plan because it feels intentional.
  • Reflection without external feedback. What follows is a self-congratulation loop that produces confident, unimproved output.
  • Prompt-only patterns. The outcome is untestable, unbudgetable behavior that vanishes under load or a model upgrade.
  • Skipping verification when a checker exists. The price is paying for reasoning where a validator would have done the job for one call.
  • Unbounded loops. You get the first cost incident that teaches the team what a step budget was for. Budgets are prerequisites, not patches.

Key Takeaways

  • Reasoning decides the next step; planning decides the route. Every component should be explicit about which question it answers.
  • A pattern is structure the trace can show: plan objects, critique passes, branch queues. If it lives only in the prompt, it is decoration.
  • ReAct is the default agent shape. Plan-and-execute fits long, parallelizable, reviewable work. Tree of Thoughts fits verifiable combinatorial search.
  • Best-of-N with an objective verifier is the production approximation of tree search, and usually the right amount of it.
  • Routing and verification cost one call each and deliver most of the practical value in this pattern family.
  • Reflection only works with external feedback: test results, checkers, user corrections. Without a signal, it is self-praise with extra steps.
  • Pattern choice is a cost decision. Make it with evaluation results on your tasks, and re-make it when models change.

FAQ

What is the ReAct pattern in AI agents?

ReAct interleaves reasoning and acting: each loop iteration produces a small reasoning step and one tool action, the result returns as an observation, and the loop continues until a final answer or budget. It outperformed reasoning-only baselines in the original 2022 paper because every step is grounded in fresh evidence.

What is plan-and-execute in AI agents?

A pattern that separates the route from the driving: a planner produces an explicit step list, an executor walks it, and a replanner revises the route when steps fail or reality diverges. It suits long-horizon tasks because the plan is inspectable, persistable, and parallelizable.

Is Tree of Thoughts used in production agents?

Rarely in full form. Full tree search multiplies model calls and needs pruning logic you must maintain. Production systems usually run the cheaper approximation: generate N candidates, verify with an objective checker, keep the best. Full search earns its keep on combinatorial problems with code-verifiable success.

Do AI agents need chain-of-thought to work?

No. The patterns in this article are orchestration structures, not prompting requirements. A loop with good tool contracts and budgets works without any specific reasoning style, and models change their internal behavior between versions anyway. Design what you can observe and control: steps, plans, checks, budgets.

What is the difference between planning and reasoning in AI agents?

Reasoning decides the next step given current observations, inside a loop iteration. Planning decides the sequence of steps for the whole run, above the iterations. Agents can have either without the other, and the architecture should say explicitly where each decision lives.

Conclusion

Planning and reasoning patterns are how you buy capability from a model, and each purchase has a price: latency, tokens, and code you must maintain. ReAct buys adaptability. Planning buys foresight. Search buys deliberation. Verification and routing, the two cheapest patterns, buy more quality per token than any of the expensive ones.

The discipline that separates working systems from impressive demos is the same in every case: the pattern must exist as structure, it must be bounded by budgets, and it must be judged by trajectory evaluation on real tasks rather than by how the transcript reads.

The pattern is what the trace shows, not what the prompt promises. Pick the cheapest one your evaluation cannot distinguish from the expensive one.

Last updated on 1 October 2026

Share this article

Leave a Reply

Your email address will not be published. Required fields are marked *