Cost Control: Token Budgets, Caching, and Model Routing

AI agent cost control: token budgets, prompt caching, model routing, and context discipline, measured as cost per successful task, not cost per call.

Executive Summary: Measure AI agent cost per successful task, not per call, because retries and failures make cheap calls expensive. This post walks through the five cost pools and the levers in order of impact: context discipline, caching, model routing, hard token budgets, precise retrieval, and retry limits, plus how to track spend in traces and alert on it.

“Use a cheaper model” is the most common cost advice for agents, and it is the least reliable. A cheap model that needs four retries, two extra steps, and a human rescue can cost more per completed task than a strong model that finishes in one pass. The invoice does not grade effort.

Agent costs are structural: input tokens reprocessed on every iteration, output tokens generated by a meter, tool calls billed per API, retries multiplying everything, and humans reviewing the residue. Structural costs respond to structural levers, not to model shopping.

This article covers those levers in order of leverage: the metric that matters, where the money goes, token budgets, caching, model routing, context discipline, and the latency-versus-cost trade that ties them together.

The Metric That Matters: Cost per Successful Task

Cost per model call flatters failing systems. If a run attempts the task four times, succeeds once, and costs one dollar per attempt, the honest number is four dollars, not one. Cost per successful task is total spend divided by tasks achieved, and it is the only figure that behaves like the business.

The metric also reorders intuitions. A prompt improvement that lifts success rate from 50 to 80 percent cuts cost per successful task by more than a third, more than a 25 percent per-call discount would, because the denominator grows. A stronger model that completes in two steps beats a cheaper one that needs six. Optimization targets the fraction, not the numerator, and the evaluation machinery that measures success is what makes the fraction computable at all: agent evaluation.

Where the Money Goes

Cost pool What drives it Controlled by
Model spend Input and output tokens, per iteration Context discipline, routing, caching
Tool and API spend Per-call pricing of external systems Caching, retrieval precision, budgets
Infrastructure Workers, state stores, queues Concurrency caps, scaling policy
Observability Telemetry volume and retention Metadata-first tracing, retention limits
Human review Approvals, escalations, corrections Approval thresholds, agent quality

Teams routinely optimize the first row and ignore the fifth. Human review is the most expensive token in the system, priced in attention, and every approval threshold set too low converts agent cost into salary cost while feeling like diligence.

The Cost Levers, Ranked by Leverage

Lever Mechanism Typical effect
Context discipline Smaller, selected windows Cut input tokens on every iteration
Exact caching Reuse for repeated calls, stable prefixes Eliminate repeat spend at the source
Model routing Cheap models for routine steps Cut per-step price on non-decision calls
Token budgets Hard caps per run and user Bound the worst case
Retrieval precision Fewer, better passages Less re-read noise per step
Retry discipline Idempotent, bounded retries Stop multiplying every other cost
Model downgrade Cheaper model everywhere Last, and only with evaluation evidence

The bottom row is deliberate. Model downgrades are a valid lever, but they belong last, applied per step with evaluation evidence, because they trade success rate, and success rate is the denominator of every other number here.

Token Budgets: Bounds Before Optimization

Budgets are the safety net that makes every other lever safe to experiment with. Three layers:

  • Per-run budgets cap the cost of one task: maximum input tokens, output tokens, and tool calls, enforced by the orchestrator before spend.
  • Per-user and per-tenant budgets cap exposure over time, enforced at the gateway, where the deployment architecture already controls access.
  • Per-task-type budgets encode business knowledge: a refund lookup and a migration analysis deserve different ceilings.

Unbounded consumption appears in OWASP’s LLM risk list for good reason: a system that spends without limit is a denial-of-wallet vulnerability, whether the spender is malicious or merely enthusiastic: OWASP Top 10 for LLM Applications.

Caching: Exact First, Semantic Rarely

Two cache types, with different trust levels:

  • Exact caching. The same tool call with the same arguments within a freshness window returns the recorded result. Same for stable prompt prefixes, where provider-side prompt caching discounts repeated input. High trust, high savings, low cleverness.
  • Semantic caching. Similar queries return similar answers. Useful for read-only, low-stakes retrieval at scale, and dangerous everywhere else, because “similar” is exactly the judgment a cache cannot make reliably for consequential tasks.

The agent-specific opportunity is larger than in chatbots: agent workloads repeat structurally. The same policy documents retrieved across a thousand support runs, the same lookups inside one long investigation, the same classification prompts across a batch. Caching the tool layer often beats caching the model layer, because a cached lookup is free of model variance entirely.

Model Routing: Pay for Decisions, Not for Chores

A multi-step run is not one task with one price tag. Routing assigns a model per step type, and the sketch is simpler than the savings suggest:

STEP_MODEL_ROUTES = {
    "compact_history": "fast-model",      # routine compression
    "classify_intent": "fast-model",      # high-volume routing
    "embed_query":     "small-embedder",  # retrieval support
    "decide_next_step": "strong-model",   # the actual agent loop
    "draft_final_answer": "strong-model",
}

def model_for(step: str) -> str:
    return STEP_MODEL_ROUTES.get(step, "strong-model")

Three rules keep routing honest. Route by step type, never by vibe: the table is explicit, reviewed, and evaluated. Measure each route with cost per successful task, because a cheap model that doubles retry count is not cheap. And keep fallback routing: when the strong model is down or throttled, degrade explicitly rather than silently switching the decision steps to a model that has never passed the suite.

What routing must not become is a replacement for the success metric. “The cheap model passes 80 percent of scenarios” is a routing decision only if the failing 20 percent does not matter, and the evaluation suite is what makes that knowable rather than assumed.

Context Discipline Is Cost Discipline

Every token in the window is paid for on every iteration. A 30,000-token context through eight model calls is 240,000 input tokens before any output is counted, which is why the single highest-leverage cost decision in any agent is what enters the window.

The levers are the ones from context window management: select by section, summarize the old trajectory, prune what state already holds, bound observations at the source. The pleasant surprise is that these are the same levers that improve quality, because noise removed from the window is both spend saved and attention recovered. The rare win-win in engineering, and the reason context discipline outranks model choice on the lever table.

Tool Economics and Retry Multiplication

Model spend is the visible cost. Tool API spend is the one that surprises teams at the end of the month: per-call pricing across a thousand runs adds up faster than tokens. Three rules keep it bounded:

  • Cache repeated lookups. Agent workloads repeat structurally. The same policy check, the same product record, the same enrichment call across runs are exact-cache candidates with freshness windows.
  • Retrieve precisely. Fewer, better passages cost less to re-read on every step that uses them. Retrieval precision is a cost lever disguised as a quality lever, and it is both.
  • Stop retry multiplication. Every retry re-pays the model call and the tool call. The reliability rules (bounded, idempotent, transient-only retries) are cost rules wearing a different badge: reliability for agents.

The unnecessary loop deserves its own line. An agent that circles pays twice for the same nothing: model calls to decide to repeat, and tool calls to repeat it. Duplicate-action detection is a cost control as much as a reliability control, and it is the cheapest one on this list.

Latency Versus Cost

The two axes trade against each other, and pretending otherwise produces surprise bills or surprise slowness:

  • Concurrency. Parallel steps cut wall-clock time and raise concurrent spend against provider rate limits. The right setting is a business decision with numbers, not a default.
  • Caching. Freshness windows trade staleness for speed and savings. Long windows on fast-changing data trade correctness, which is not on the table.
  • Strong versus cheap models. The strong model finishing in one pass often wins both axes, latency and total cost, over a cheap model needing several attempts. The evaluation suite is what turns this from an anecdote into a per-task fact.

Set the latency requirement first, then optimize cost under it. Optimizing cost without a latency constraint quietly optimizes user patience instead.

The Arithmetic, Worked

Cost control starts with being able to compute the number you are trying to move. Take one task type and instrument a week of runs, then fill in four values per run: input tokens, output tokens, tool calls at their per-call prices, and retries. Suppose a representative run carries about 12,000 input tokens across four iterations, 1,500 output tokens, and two paid tool calls. Plug in your provider’s published rates, and per-run model cost is input tokens times the input rate plus output tokens times the output rate, plus tool spend, with retries multiplying everything. The formula is boring; running it before every optimization is the discipline.

The levers then read directly off the arithmetic. Cutting context 40 percent cuts the dominant term 40 percent on every call. Raising success rate from 80 to 90 percent cuts cost per successful task by about 11 percent, the ratio 0.8 over 0.9, before any model change. And a model 20 percent cheaper per call that drops success from 80 to 60 percent costs about 7 percent more per success, because 0.8 divided by 0.6 exceeds 1 divided by 0.8: the per-call discount is smaller than the success-rate loss. Every lever in the table above can be checked with this arithmetic before it is believed.

Spend Governance

Agent spend deserves the same governance cloud spend learned the hard way:

  • Tag spend by agent version and task type, the way cloud costs are tagged by service and team, so regressions attribute within hours.
  • Alert on distributions, not just totals. A cost-per-task p95 that doubles is a regression even while the average looks flat.
  • Make budgets code, enforced at the gateway, with showback per team so the people starting runs see what runs cost.
  • Re-run the lever table quarterly. Rates, models, and caching options change; last quarter’s optimal routing table is this quarter’s assumption.

Cost discipline is familiar ground for me: I led an infrastructure modernization that cut cloud operating costs by over 80 percent through platform rationalization, architectural redesign, and automated scaling. Token spend responds to the same arithmetic-first treatment: measure the fraction, move the biggest term, verify the move. The backend side of that playbook is in cost optimization for backend systems, and the per-call basics for LLM apps are in the LLM app production checklist.

How Real Systems Do This

  • Cost per successful task sits next to success rate on the main dashboard. The two numbers together answer the only business question: does this system work, and what does working cost.
  • Routing tables are reviewed like configs. Every route change runs the evaluation suite, because routing is a behavior change.
  • Exact caching lives at the tool gateway, with explicit freshness windows per source, and provider prompt caching handles stable prefixes.
  • Budgets and alerts are deployment features, enforced at the gateway per user and tenant, with anomaly alerts on spend that deviate from the per-task distribution.
  • Cost is observable per step. Traces carry token and cost attributes, and platforms like Langfuse surface cost and latency analysis directly, so optimization targets steps rather than vibes.

Decision Framework

Run this sequence for any agent with real volume:

  1. What does a successful task cost today? Baseline first: total spend divided by passed tasks, from trace data. No lever is rankable without this number.
  2. Which pool dominates? Model, tools, infrastructure, or human review. Optimize the biggest pool first, and remember the salary pool is real.
  3. Is the context selected? If the window accumulates rather than selects, context discipline is the biggest move available.
  4. What repeats? Cache the repeated lookups and stable prefixes with freshness windows. This is found money.
  5. Which steps are chores? Compaction, classification, formatting: route them to cheap models. Decisions stay on the strong one.
  6. Where are the caps? Per-run, per-user, per-task-type budgets, enforced before spend, with alerts on the anomalies.
  7. What does latency require? Fix the constraint, then optimize cost under it, including the concurrency math.

When NOT to Use This

  • Low-volume prototypes. Optimization without traffic is guesswork with spreadsheets. Install budgets, then revisit when volume exists.
  • When quality is the product and the spend is approved. Cost discipline still means budgets and anomaly alerts, but routing gymnastics and cache complexity are not owed to a system whose economics already work.
  • When success cannot be measured. Cost per successful task is undefined without a success definition. Fix evaluation first, or every optimization optimizes an unmeasurable.
  • Semantic caching on consequential tasks. “Probably the same answer” is a savings technique for read-only, low-stakes queries, not for anything with side effects or users’ money in it.

Common Mistakes

  • Downgrading the model everywhere. Where it lands: success rate falls, retries rise, and cost per successful task goes up while the per-call dashboard goes down.
  • Dashboards in cost per call. The consequence: every optimization that hurts success looks free, and the fraction that matters is never computed.
  • No per-user caps. What follows: one enthusiastic user converts a cost program into an incident.
  • Caching without freshness policy. The price: stale answers delivered at a discount, which is the worst possible pricing.
  • Human review costs off the books. In practice: approval thresholds set by feel, converting model savings into salary at an unfavorable exchange rate.
  • Retries without idempotency. The result: duplicated side effects, which cost money twice and then cost money a third time in cleanup.

Key Takeaways

  • Cost per successful task is the metric that behaves like the business. Cost per call flatters failing systems.
  • Five cost pools: model, tool APIs, infrastructure, observability, human review. The last one is the most expensive and the least tracked.
  • Leverage order: context discipline, exact caching, model routing, token budgets, retrieval precision, retry discipline. Model downgrade comes last, with evaluation evidence.
  • Context is the strongest lever because input tokens are reprocessed every iteration, and shrinking the window improves quality at the same time.
  • Routing pays for decisions and saves on chores, but every route is a behavior change that runs the suite.
  • Budgets bound the worst case: per run, per user, per task type, enforced before spend. Unbounded consumption is a named security risk.
  • Fix the latency requirement first, then optimize cost under it.

FAQ

How much does it cost to run an AI agent?

It is a distribution, not a number: model spend scales with steps and context, tool APIs bill per call, retries multiply both, and human review prices in attention. The useful expression is cost per successful task, computed from trace data, tracked per task type and agent version.

How do you reduce AI agent costs?

In leverage order: shrink and select the context, cache repeated calls and stable prefixes exactly, route routine steps to cheaper models, enforce token budgets per run and user, retrieve precisely, and discipline retries. All six require the success metric to verify, because savings that hurt success rate are not savings.

Is switching to a cheaper model a good cost strategy?

Only per step and only with evidence. A cheap model that doubles retry counts or halves success rate usually costs more per successful task than a strong model that finishes in one pass. Route chores to cheap models, keep decisions on strong ones, and let the evaluation suite arbitrate.

What is model routing in AI agents?

Assigning a model per step type instead of one model for the whole run: classification and compaction go to fast, cheap models, while decision steps and final answers go to strong models. The routing table is explicit, versioned, and every change runs the evaluation suite, because routing changes behavior.

Does prompt caching actually reduce agent costs?

Yes, where the prefix is stable across calls, which agent loops often are: repeated system instructions and stable context heads. It discounts the bill rather than the attention budget, so treat it as a cost lever, not a substitute for context selection.

Conclusion

Agent cost control is not a shopping problem. It is a systems problem with five pools, one honest metric, and a ranked set of levers, most of which improve quality while they reduce spend. The teams that win on cost are the teams that measure the fraction, select the context, cache what repeats, route the chores, and bound everything.

The failure pattern is equally consistent: per-call dashboards, model downgrades without evaluation, no caps, and human review hiding off the books. Every one of those looks like savings in a meeting and looks different on an invoice.

The cheapest agent is the one that finishes. Spend on the steps that decide, save on everything else, and always divide by success.

Last updated on 4 October 2026

Share this article

Leave a Reply

Your email address will not be published. Required fields are marked *