Reliability for Agents: Retries, Timeouts, Fallbacks, Circuit Breakers
Reliability for agents: the failure taxonomy plus retries, timeouts, fallbacks, circuit breakers, idempotency, and bounded execution that keep runs alive.
An API fails one way: it returns an error, and you handle it. An agent fails eleven ways, and the ways interact. The tool times out, so the loop retries, but the first call actually went through, so the retry duplicates the side effect, and the model, seeing two confusing observations, calls the same tool again.
Agent reliability is not one technique. It is a small toolkit (timeouts, bounded retries with backoff, fallbacks, circuit breakers, and idempotency) applied to a specific failure taxonomy, plus one thing ordinary services never needed: bounded execution, because an agent can fail by succeeding indefinitely.
This article covers the taxonomy, the toolkit, and the wiring: where each technique belongs in the loop, what it costs, and where teams get it wrong.
How Agents Fail: The Taxonomy
| Failure | Layer | First symptom | Primary defense |
|---|---|---|---|
| Malformed tool arguments | Decision | Schema validation error | Validate before execution, return structured error |
| Invalid tool response | Observation | Model confusion next step | Result validation and shaping |
| Model refusal or empty output | Decision | Run stalls or ends thin | Retries with variation, fallback prompt |
| Hallucinated result | Decision | Confidently wrong action | Verification on critical outputs |
| Infinite loop | Loop | Step count climbs, output stalls | Step budget, duplicate detection |
| Repeated action | Execution | Same tool call twice | Duplicate-action guard, idempotency keys |
| Context overflow | Assembly | Provider error mid-run | Context budgets, bounded observations |
| Stale memory | Assembly | Confident outdated claims | Freshness rules, timestamps in context |
| Transient provider failure | Model call | Timeouts, 5xx, throttling | Bounded retry with backoff |
| Partial success | Run | Some steps done, run ends | Checkpoints, resume, explicit status |
| Duplicated side effect | Execution | Two refunds, two emails | Idempotency keys, at-most-once writes |
Read the table as a wiring diagram. Each row has a defense, each defense has a place in the loop, and the place matters: validation belongs before execution, budgets belong at the top of the iteration, and duplicate guards belong at the tool gateway, the choke point where the tool layer meets the loop.
Timeouts: Every Call Gets One
Three timeout layers, each protecting a different thing:
- Tool call timeout protects the loop from a hung dependency. Set per tool, informed by its observed latency profile, with percentiles rather than averages.
- Model call timeout protects the run from a stalled provider. Generous, because long generations are legitimate, but finite.
- Run wall-clock timeout protects the platform and the user. Nothing waits forever, including background jobs.
The subtle rule is about writes. A timeout does not mean the call failed; it means the outcome is unknown. For a read, unknown is safe to retry. For a write, unknown means the effect may have happened, and blind retries are how duplicate refunds are born. The timeout policy for every write tool must therefore be: stop, check what happened through a status or idempotency mechanism, then act.
Retries: Bounded, Backed-off, Idempotent-Only
Retries are the most misapplied technique in agent systems, because the default instinct (retry failed calls) is safe in ordinary services and dangerous here. Three rules make retries safe:
- Retry only transient failures. Timeouts on reads, provider 5xx, throttling. Never retry schema violations, authorization denials, or model refusals: those will fail identically, and the retry only delays the truth.
- Bounded attempts with exponential backoff and jitter. Three attempts with growing, randomized delays is a default that survives contact with real dependency behavior. Unbounded retries are loops with a respectable name.
- Never automatically retry non-idempotent writes. If the write has no idempotency key, the failure observation goes to the model or a human, not back to the wire.
The example is illustrative; the structure is the point. Failed retries end as structured error observations, per the failure-handling rules of the tool layer:
import random
import time
from typing import Callable
class TransientError(Exception):
"""Timeouts, 5xx responses, throttling: failures that may recover."""
def call_with_retry(func: Callable[[], dict], attempts: int = 3,
base_delay_s: float = 0.5) -> dict:
"""Retry transient failures with exponential backoff and jitter."""
last_error = "no attempt made"
for attempt in range(1, attempts + 1):
try:
return func()
except TransientError as exc:
last_error = str(exc)
if attempt == attempts:
break
delay = base_delay_s * (2 ** (attempt - 1))
time.sleep(delay + random.uniform(0, 0.1))
return {"error": f"failed after {attempts} attempts: {last_error}"}
Fallbacks and Graceful Degradation
When retries are exhausted, the question is what the run does next, and the honest answer is a ladder:
- Alternate route. A different tool, source, or model that can serve the same step.
- Reduced capability. Complete the task partially, and say so: “I could not verify the shipping status, but here is the order summary.”
- Honest failure. Stop, report what was completed, checkpoint, and escalate.
The forbidden option is the fourth: faking success. An agent that invents a plausible answer when its evidence failed is worse than a crashed agent, because the failure is undetectable at the interface and expensive at the consequence. Degradation must be visible in the output and in the trace, per the span taxonomy of tracing agent runs.
Circuit Breakers: For the Model That Keeps Believing
Circuit breakers exist in ordinary services, but agents have a special need for them. A human engineer stops using a dependency that has failed five times in a row. A model has no such social instinct: if the tool contract says it works, the model will keep calling it, every run, forever, burning budget and patience.
The mechanics are standard: after a failure threshold within a window, the breaker opens and calls fail fast, immediately, without touching the dependency. After a cooldown, half-open probes test recovery with a single allowed call. The agent-side addition is what the open breaker tells the model: a structured observation, “this tool is temporarily unavailable, use an alternative or wait,” which lets the loop route around the outage instead of hammering it.
Idempotency: Duplicate Effects Are the Agent Tax
Agents act, and distributed systems retry. The intersection is duplicated side effects, and it deserves the most engineering care of anything in this article because it is the failure mode that reaches users directly.
The fix is mechanical: every write tool accepts an idempotency key, derived from the run and the step; the executing system records the key with the outcome; a retried call with the same key returns the recorded outcome instead of executing again. With that in place, the platform can deliver at-least-once semantics safely, and the deployment can redeliver jobs and resume runs without fear. The state machinery that makes resume safe is agent state persistence.
Bounded Execution: Failing by Succeeding Indefinitely
Ordinary services stop when their work stops. An agent can keep working productively forever (checking one more source, refining one more draft), and the bill follows the effort. Bounded execution is therefore a reliability requirement, not a cost optimization:
- Step budget. Maximum iterations per run, enforced before each model call.
- Token budget. Maximum total input and output tokens, enforced at assembly time.
- Wall-clock budget. Maximum run duration, enforced by the platform, not the model’s sense of time.
- Duplicate-action detection. The same tool with the same arguments twice in a row is a stuck loop wearing progress clothing. Block it, and tell the model why.
Unbounded consumption is listed in OWASP’s LLM risk set for exactly this reason: a system that spends without limit is a security exposure as much as a cost problem: OWASP Top 10 for LLM Applications. Durable-execution runtimes recognize the same concern with built-in recursion limits and fault-tolerance features; LangGraph, for example, documents recursion limits and fault tolerance as first-class capabilities: LangGraph documentation.
How Real Systems Do This
- Budgets run first. The loop checks step, token, and time budgets before spending, every iteration, and the check is code, never a prompt request.
- Errors are observations. Tool failures reach the model as structured data with recovery hints, which converts most reliability events into successful recoveries rather than dead runs.
- Idempotency keys are mandatory for writes. No write tool ships without one, and the gateway rejects keyless writes in review, not in production.
- Breakers sit at the tool gateway. Chronic failure opens the circuit for everyone, with a structured observation routed to the affected agents.
- Partial success is explicit. Budget-exhausted runs return what was completed, checkpointed, with status “partial”, rather than pretending completion or silently truncating.
- Reliability metrics come from traces. Retry counts, breaker events, and timeout rates are span aggregates, watched as trends per tool and per model version.
Decision Framework
For every agent going to production, answer in order:
- What is the timeout for each tool? From observed latency percentiles, per tool, not one global number.
- What is retryable? Enumerate transient failures per dependency. Everything else fails fast with an observation.
- What is the retry policy? Bounded attempts, exponential backoff with jitter, and idempotency required for writes.
- What is the degradation ladder? Alternative, then partial with disclosure, then honest failure. Fake success is not an option, so it must not be one in code either.
- Which tools get circuit breakers? The flaky ones first, then all external dependencies.
- What are the budgets? Steps, tokens, wall clock, and duplicate detection, all enforced in the orchestrator.
- How are failures visible? Every retry, breaker event, and budget trip appears in the trace and in the failure metrics.
When NOT to Use This
- Prototypes. A demo agent with a step budget and a timeout on each model call is enough. Full reliability machinery on an unstable loop is polishing a prototype.
- When failure is cheap and visible. Internal tools with read-only effects can fail loudly and be rerun by the user. Spend engineering where consequences live.
- As a substitute for good contracts. Retries around a badly described tool convert fast failures into slow ones. Fix the contract first.
- When it masks a design error. If the agent needs constant retries to complete ordinary tasks, the problem is the design, and the toolkit is an anesthetic.
Common Mistakes
- Retrying non-idempotent writes. What you get: duplicate refunds and duplicate emails, delivered politely by the reliability system.
- Global timeouts. Where it lands: fast tools over-waited, slow tools killed early, and both blamed on the model.
- Unbounded retries. The consequence: a dependency outage multiplied into a cost and latency incident.
- Errors as exceptions. What follows: runs that die instead of adapt, wasting the loop’s core advantage.
- No breaker on flaky tools. The price: a model that keeps calling a dead tool every run, because the contract says it works.
- Budgets as prompts. In practice: an “instruction” to be concise or stop after N steps, ignored precisely when it matters, because the model was never the place to enforce it.
Key Takeaways
- Agents fail in two layers: ordinary distributed-system failures and agent-specific ones: hallucinated arguments, loops, duplicate actions, context overflow, partial success. Each failure has a named defense and a place in the loop.
- Every call gets a timeout: per tool, per model call, per run. For writes, a timeout means unknown outcome, and unknown is not safe to retry blindly.
- Retries are bounded, backed off with jitter, and idempotent-only. Never retry schema violations, authorization denials, or refusals.
- Degrade honestly: alternate route, partial result with disclosure, then explicit failure. Faking success is the one forbidden option.
- Circuit breakers matter more for agents than for services, because the model will keep calling a broken tool as long as the contract says it exists.
- Idempotency keys on every write make at-least-once delivery and run resumption safe. Duplicate effects are the failure mode users meet first.
- Bounded execution (steps, tokens, wall clock, duplicate detection) is a reliability and security requirement, enforced in code before spend.
FAQ
How do you make AI agents reliable?
Map the failure taxonomy first: malformed arguments, invalid responses, transient provider errors, loops, duplicates, context overflow, partial success. Then wire the toolkit to each failure: timeouts on every call, bounded idempotent-only retries, graceful degradation ladders, circuit breakers on flaky tools, idempotency keys on writes, and hard budgets in the orchestrator.
Should you retry failed AI agent tool calls?
Only transient failures on safe operations: timeouts on reads, provider 5xx, throttling. Never retry schema violations or authorization denials, which fail identically, and never automatically retry non-idempotent writes, where a timeout means the outcome is unknown and the retry may duplicate the effect.
What is a circuit breaker in agent systems?
A guard at the tool gateway that stops calling a dependency after a failure threshold within a window, fails fast with a structured observation, and probes recovery after a cooldown. Agents need them especially because the model will keep calling a broken tool as long as its contract suggests it works.
Why are duplicate side effects an agent reliability problem?
Because agents act and distributed systems retry. A timeout or worker crash leaves the outcome unknown, and a naive retry executes the action twice. Idempotency keys derived from the run make retries safe by returning the recorded outcome instead of executing again.
How do you stop runaway agent loops?
With mechanical bounds enforced before spend: step and token budgets checked every iteration, a wall-clock timeout, and duplicate-action detection that blocks the same call with the same arguments twice in a row. All of it lives in the orchestrator, never in the prompt.
Conclusion
Reliability engineering for agents is ordinary engineering with one added law: the system can fail by continuing. Every technique in this article exists to keep a nondeterministic component inside a bounded, observable, recoverable envelope, and the envelope is the product.
The teams that operate agents well do not have better models. They have better defaults: every call timed out, every write idempotent, every failure an observation, every loop budgeted, and every degradation honest. None of it is glamorous, and all of it compounds.
The rule of thumb: retry what is safe, degrade what is failing, and bound everything that can spend.
Last updated on 5 October 2026
