How to Evaluate AI Agents: Task Success, Trajectory, and Cost

Evaluate AI agents beyond answer quality: task success, tool and argument correctness, trajectory quality, cost per successful task, latency, and recovery.

Executive Summary: Evaluate AI agents on three layers: task success, trajectory quality, and economics such as cost per successful task and latency. This post shows how to build a versioned scenario suite from real traces, when to use objective checkers versus a validated LLM judge, and why every prompt, tool, or model change should rerun the full suite.

A model answers, so you grade the answer. An agent acts, so the answer is the least of what you must grade. The refund was correct, but it went to the wrong account, used three times the necessary steps, and would have looped forever without the step budget. Answer quality saw none of that.

Agent evaluation is the difference between shipping a demo and operating a system. It answers three questions continuously: does the agent complete tasks, does it take acceptable paths to do so, and what does each success cost?

For the single-call case, how to test and evaluate LLM apps in Python covers the fundamentals this article extends to multi-step runs. This article covers the evaluation stack: the metric set, how to build scenario suites, grading approaches from rules to judges to humans, trajectory review, cost per successful task, and the regression discipline that keeps upgrades from becoming regressions. The observability companion, tracing, is covered in tracing agent runs; evaluation asks the questions, tracing supplies the evidence.

Why Final-Answer Evaluation Fails for Agents

Grading only the final output fails in two directions. It passes bad agents: a correct answer reached through a hallucinated tool call, a lucky guess, or a policy violation ships as a success. And it fails good agents: a run that escalated correctly at the right moment scores lower than one that blundered through to an answer.

The root cause is that the outcome is underdetermined by the answer. The same answer can come from a two-step efficient run or a nine-step loop that happened to converge. For chatbots, where the product is the text, answer quality is the product. For agents, where the product is the run, the answer is one artifact among many, and the least informative one about how the system will behave next week.

The Metric Set

Metric What it measures How to check it
Task success rate Runs that achieve the defined outcome Objective checker per scenario: state change, API call, output schema
Tool-selection correctness Right tool for the step Expected tool per scenario step, checked against the trajectory
Tool argument correctness Valid, safe arguments Schema validation plus expected-values checks on critical arguments
Trajectory quality Sensible order, no loops, no dead weight Step count vs reference, duplicate-action detection, loop detection
Policy compliance Permission and approval rules respected Assert no out-of-policy tool calls; approvals requested when required
Cost per successful task Tokens, tool calls, retries, divided by successes Sum run costs, divide by passed runs, from cost control data
Latency Wall-clock per task and per step Timestamped traces; percentiles, not averages
Recovery behavior Response to injected tool failures Chaos scenarios: fail a tool, assert correct retry or escalation
Escalation rate How often humans take over Escalation events per task, tracked over time

Not every metric belongs in every suite. Task success and cost per successful task are non-negotiable. The rest enter when the corresponding risk is real: policy compliance when the agent touches money or data, recovery when it depends on flaky tools, escalation rate when humans are in the loop.

Building the Scenario Suite

The suite is the heart of evaluation, and it has the same status as production code: versioned, reviewed, and grown deliberately.

  • Start from real traces. Every production incident and every surprising success is a scenario candidate. Suites built only from imagined happy paths measure imagination.
  • Define success as a check, not a description. “Refunds the eligible amount to the correct account and cites the policy” is an assertion; “handles the request well” is not.
  • Cover the distribution, not just the average. Include the rare-but-critical cases: the refund above the threshold, the angry customer, the tool that returns empty. Agents fail at the tails.
  • Version scenarios with the system. When the agent’s capabilities change, the suite changes, and old results stay comparable through dataset versioning.
  • Keep chaos scenarios. A handful of scenarios with deliberately failing tools measures recovery, which is where production agents actually differ from demos.

Grading: Rules, Judges, and Humans

Three grading mechanisms, in order of preference:

  1. Objective checkers. Database state, API call assertions, output schema validation, expected tool sequences. Deterministic, cheap, and trustworthy. Anything expressible as a check should be.
  2. LLM-as-judge. For soft criteria: tone, completeness, “does this reply address the question.” Useful but not free: judges are models, with their own biases and drift. Validate a judge against human labels on a sample before trusting it, and re-validate when the judge model changes. Platforms like LangSmith and Langfuse both support judge-based evaluations, datasets, and annotation queues for exactly this workflow: LangSmith, Langfuse.
  3. Human review. For the residual: novel failure patterns, judgment calls, and periodic sampling of production runs to find what the suite is missing. Expensive, so spend it on discovery, not routine grading.

A common failure is inverting the order, sending everything through a judge because it is easier to set up. The result is an evaluation suite that is itself nondeterministic, grading a nondeterministic system with a nondeterministic grader, and producing numbers nobody trusts. Determinism is a property worth paying for.

Trajectory Review: Reading the Path

Trajectory evaluation inspects the run, not just its ending. Every trajectory failure has a shape, and naming the shape is most of the fix:

Failure shape What the trace shows Typical root cause
Wrong tool Call unrelated to the step’s need Tool description or contract ambiguity
Right tool, wrong arguments Schema-valid but wrong values Missing context, poor argument examples
Right call, wrong conclusion Correct result, misread reasoning Observation too noisy or too large
Redundant steps Repeated calls, duplicate fetches No working state, forgotten results
Looping Same call cycling until budget No duplicate detection, recovery logic absent
Premature answer Answer before evidence gathered Prompt pressure toward speed, weak stopping rules

This is why the metric set and the pattern language from planning and reasoning patterns connect: trajectory failures are usually pattern or contract failures, not model intelligence failures, and they are fixed in contracts and context, not by swapping models.

Cost per Successful Task

Per-call cost is the wrong denominator for agents, because a run that loops four times and fails costs four calls for zero value. Cost per successful task divides everything (model tokens, tool API spend, retries, and infrastructure) by passed runs.

That number reframes common decisions. A strong model that completes tasks in two steps can beat a cheap model that needs six. A prompt improvement that lifts success rate from 60 to 90 percent cuts cost per successful task by a third, more than a model downgrade that saves 20 percent per call, because the denominator grows. Track the full cost model, including human review time, in the framework from cost control.

A Minimal Evaluation Harness

The harness below is illustrative and deliberately plain: scenarios with objective checkers and budgets, run against the agent, reported as pass rates and over-budget counts. Platform tools like LangSmith and Langfuse give you the same structure with datasets, dashboards, and trend lines; the logic is identical.

from dataclasses import dataclass
from typing import Callable

@dataclass(frozen=True)
class Scenario:
    name: str
    task: str
    check: Callable[[dict], tuple[bool, str]]   # (passed, reason)
    max_cost_usd: float = 0.50
    max_latency_s: float = 30.0

def run_suite(agent, scenarios: list[Scenario]) -> dict:
    """Minimal harness: run scenarios, grade outcomes. Illustrative."""
    results = []
    for scenario in scenarios:
        outcome = agent.run(scenario.task)
        passed, reason = scenario.check(outcome)
        within_budget = (outcome.cost_usd <= scenario.max_cost_usd
                         and outcome.latency_s <= scenario.max_latency_s)
        results.append({
            "scenario": scenario.name,
            "passed": passed,
            "reason": reason,
            "cost_usd": outcome.cost_usd,
            "latency_s": outcome.latency_s,
            "within_budget": within_budget,
        })
    passed_count = sum(r["passed"] for r in results)
    return {
        "total": len(results),
        "passed": passed_count,
        "over_budget": sum(not r["within_budget"] for r in results),
        "details": results,
    }

The design commitments that matter: success is a function you wrote, not an opinion; budgets are per scenario because tasks differ; and the report separates correctness from economics, so a change that improves one and damages the other is visible instead of averaged away.

Regression Discipline

An evaluation suite only protects you if it runs when things change, and agents change through channels that traditional CI ignores:

  • Prompt changes are releases. They alter behavior as surely as code.
  • Tool changes are releases. A modified tool description or result shape changes trajectories.
  • Model upgrades are releases. New model versions shift tool selection, argument formatting, and stopping behavior in ways no changelog predicts.
  • Retrieval corpus changes are releases for agents that depend on retrieval.

The practical rule: run the suite on every pull request that touches any of the four, with thresholds that block merges, and archive the results per version so trends are visible. Model upgrades deserve a full suite run on a schedule, because providers change behavior without asking your permission.

How Real Systems Do This

  • Success definitions on my flagship systems are objective by construction. My engineering intelligence agent defines success as leadership-ready insights grounded in delivery telemetry, not as “good analysis.” If your suite cannot express the success definition, the product will invent a vaguer one.
  • Suites run in CI, on real triggers. Prompt diffs and tool-contract diffs invoke evaluation the way unit tests invoke themselves.
  • Production traces feed the suite. Failed runs become new scenarios within the week, which is why mature suites look like the product’s actual failure distribution.
  • Online evaluation runs on live traces. Sampled production runs get graded continuously, with alerting on success-rate drops, cost spikes, and policy violations, using the same platforms used for tracing.
  • Judges are validated, then trusted with scope. Soft-criteria grading runs through LLM-as-judge after validation against human labels, and judge outputs feed dashboards rather than merge decisions where determinism matters.
  • Cost per successful task is the headline number, not token counts, because it is the number that behaves like the business.

Decision Framework

Before writing evaluation code, decide these in order:

  1. Can each task’s success be defined as a check? If no, define the task better before evaluating the agent. Unmeasurable tasks produce unmeasurable agents.
  2. Which metrics match the risks? Success and cost always. Policy compliance for money and data. Recovery for flaky dependencies. Escalation rate for human-in-the-loop systems.
  3. Where does grading determinism matter? Merge-blocking checks should be objective. Judges and humans grade discovery and soft criteria.
  4. What triggers the suite? Prompt, tool, model, and corpus changes. If the trigger list is shorter than that, the suite is decoration.
  5. What are the thresholds? Pass rates that block, budgets that warn, trends that page. Numbers written before the first regression argument.
  6. How do failures return to the suite? Every production incident ends as a scenario, or the same incident returns.

When NOT to Use This

  • Prototypes and exploratory demos. Full evaluation on a system whose shape changes weekly is waste. Keep the discipline minimal: a handful of manual test tasks and a step budget.
  • Tasks without objective success signals. If success cannot be defined, evaluation will measure judge preferences instead of product quality. Fix the task definition first.
  • Low-stakes internal tools. A single-user summarizer with read-only inputs may not justify a suite. Weigh the blast radius honestly, and revisit when it grows.
  • When it substitutes for observability. Evaluation measures the suite’s distribution; tracing shows what production actually does. Skipping tracing because “we have evals” leaves you blind to the divergence between the two.

Common Mistakes

  • Grading only final answers. The damage: policy violations, lucky guesses, and looping runs all ship as green.
  • Judging everything with a judge. The fallout: nondeterministic grades of a nondeterministic system, and dashboards nobody defends in a design review.
  • Suites of imagined happy paths. The cost: high scores in CI and a completely different failure distribution in production.
  • Cost per call as the headline. What you get: model downgrades that look like savings while cost per successful task rises through retries.
  • Evals that never run. Where it lands: a beautiful dataset and a system that regresses on every quiet Friday deploy.
  • No feedback loop from production. The consequence: the suite ages into fiction while the real failure modes evolve without opposition.

Key Takeaways

  • Evaluate three layers: task success, trajectory quality, and economics. The final answer is the least informative artifact of an agent run.
  • Build a versioned scenario suite from real traces, with success defined as checks, tails included, and chaos scenarios for recovery.
  • Grade with objective checkers first, validated LLM-as-judge second, humans for discovery. Determinism is worth paying for.
  • Read trajectories as failure shapes: wrong tool, wrong arguments, wrong conclusion, redundant steps, loops, premature answers. Each shape points to a contract or context fix.
  • Track cost per successful task, not cost per call. Retries make cheap expensive.
  • Prompt, tool, model, and corpus changes are releases. Run the suite on all four, with thresholds that block.

FAQ

How do you evaluate an AI agent?

With a versioned scenario suite graded on three layers: task success via objective checkers, trajectory quality via expected tools and step analysis against recorded traces, and economics via cost per successful task and latency percentiles. Run it on every prompt, tool, model, or corpus change.

What metrics matter for AI agents?

Task success rate, tool-selection correctness, tool argument correctness, trajectory quality, policy compliance, cost per successful task, latency, recovery behavior on injected failures, and human escalation rate. Task success and cost are non-negotiable; the rest enter when their risks are real.

What is trajectory evaluation?

Evaluating the path, not just the outcome: whether the agent called the right tools with correct arguments in a sensible order, without loops, duplicates, or policy violations. It catches runs that reach correct answers through dangerous paths, which final-answer grading passes silently.

Should you use LLM-as-judge for agent evaluation?

For soft criteria, yes, within limits. Validate the judge against human labels first, re-validate when the judge model changes, and never let it block merges where a deterministic checker could. Judges drift like any model; your merge gates should not.

What is cost per successful task?

Total spend, model tokens, tool API calls, retries, infrastructure, and human review, divided by the number of runs that achieved the defined outcome. It is the honest denominator for agents because failures cost money too, and retry-heavy cheap models often lose to efficient strong ones.

Conclusion

Evaluation is what turns an agent from a demonstration into a system. The demo measures whether the idea can work. The suite measures whether the system does work, on the distribution that matters, at a cost the business can afford, and after every change that could break it.

None of it is exotic. It is ordinary software discipline, extended to cover prompts, tool contracts, and model versions as first-class release artifacts. The teams that internalize this ship upgrades confidently and sleep through provider model changes. The teams that skip it discover every regression through users, at the worst possible price.

Judge the task, inspect the trajectory, and always divide by success.

Last updated on 7 October 2026

Share this article

Leave a Reply

Your email address will not be published. Required fields are marked *