Deploying AI Agents: Architecture and Patterns

Deploying AI agents in production: execution boundary, gateway, orchestrator, tool gateway, state stores, queues, workers, and sync vs async patterns.

Executive Summary: Deploying an AI agent means running a small distributed system: an API gateway, an orchestrator with budgets, a tool gateway, a checkpointed state store, a queue with workers, and tracing. This post explains when a run can live inside a request versus a background job, how to scale workers on queue depth, where secrets belong, and how to prevent duplicate effects and runaway cost.

A prototype agent is a function you call. A production agent is a distributed system with a gateway in front, workers behind, state on the side, and a bill attached. The distance between those two shapes is where most agent projects stall.

The good news: nothing in the production architecture is exotic. It is the standard machinery of service deployment, plus three agent-specific constraints: runs are long and metered, decisions are nondeterministic, and actions need permission gates.

This article walks the full execution boundary: the components and their responsibilities, the synchronous-versus-asynchronous decision, queues and workers, state and resumability, scaling, and the security boundaries that keep the whole thing defensible.

The Execution Boundary

Draw the boundary before writing any code. Every component below exists for a reason you can name, and the diagram is the architecture review artifact:

Client
  |
  v
API / Gateway
  |
  v
Agent Orchestrator ---------------- Observability
  |
  +--> Model Provider
  |
  +--> Tool Gateway
  |      +--> Internal API
  |      +--> Database
  |      +--> SaaS API
  |
  +--> Memory / State Store
  |
  +--> Retrieval Layer
  |
  +--> Queue / Worker (long runs)
  |
  +--> Observability
Component Responsibility Agent-specific twist
API gateway Auth, request validation, rate limits Per-user cost caps, not just request caps
Orchestrator Runs the loop, enforces budgets and stopping conditions Nondeterministic decisions require traces, not just logs
Model provider Decisions Latency and cost variance; provider-side failures need bounded retries
Tool gateway Authorization, validation, rate limiting, audit The security choke point for every action, from permission design
State store Working state, checkpoints Resume after crash without repeating side effects
Queue and workers Asynchronous long runs Graceful shutdown mid-run; checkpoint on every step
Retrieval layer Evidence for decisions Untrusted content entering the loop
Observability Traces, metrics, alerts Content privacy rules, per tracing agent runs and OpenTelemetry

Synchronous or Asynchronous: The First Decision

Execution mode is the highest-leverage deployment decision, because it shapes everything downstream: state, scaling, failure handling, and the user experience.

Dimension Synchronous (request-scoped) Asynchronous (job)
Run duration Seconds Minutes to hours
State In-memory, lost on timeout Checkpointed, resumable
Failure recovery Client retries the whole run Worker resumes from last checkpoint
Scaling Request-based Queue depth and worker pool
User experience Blocking wait Status polling or completion event
Cost control Request timeout as the ceiling Step and token budgets as the ceiling

Two rules keep the decision honest. First, any run that can call more than two or three tools probably belongs in a queue, because request timeouts will kill it at the worst moment (mid-investigation, with state that dies with the request). The same ceilings bite harder on functions platforms, as covered in serverless limits. Second, idempotency is the entry fee for asynchronous execution: a worker crash and a queue redelivery must not produce a second refund. Every write tool takes an idempotency key, and every run resumable design follows checkpointing and resumable workflows.

The common hybrid acknowledges reality: the API accepts the task, validates it, enforces caps, and returns a run id immediately, and the work happens in the background. The client polls or subscribes. Users forgive waiting when they can see progress; they do not forgive vanished work.

The Request Path: Gateway Discipline

Everything the client touches is ordinary, hardened API engineering, with agent additions:

  • Authentication and authorization at the edge, before any model spend happens.
  • Input validation with schemas, because malformed tasks produce expensive garbage.
  • Rate limits per user and per tenant, for requests and for runs.
  • Cost caps per user, day, and task type. A user who can start ten thousand runs can start ten thousand bills; the gateway is where that stops.
  • Run id assignment at acceptance, so the client can poll status from the first second.

The agent-specific insight: gateway limits are cost controls, not just load controls. An agent platform without per-user spend ceilings is a denial-of-wallet vulnerability waiting for one ambitious user.

Orchestrator, Workers, and Scaling

The orchestrator is the loop from the anatomy of an agent, deployed as a service. Two properties dominate its production behavior:

  • Concurrency caps per worker. An agent run is token-bound and network-bound, not CPU-bound. A worker that handles hundreds of concurrent API requests will strangle on a handful of concurrent runs making model calls. Tune worker concurrency on provider rate limits and budget enforcement, not on CPU.
  • Graceful shutdown. Workers get drained mid-run by deploys, so shutdown must pause at a step boundary, checkpoint, and let the queue redeliver. Workers that die mid-step without checkpoints turn every deploy into lost work.

Scaling follows queue depth, with one agent-specific caution: autoscaling workers linearly with queue depth multiplies provider spend linearly with incident size. Cap the worker pool, queue the overflow, and make cost part of the scaling policy. On Kubernetes, the standard primitives apply: horizontal scaling on queue metrics, liveness and readiness probes, and pod disruption budgets for drains; the Kubernetes documentation covers the mechanics, and agent workloads mostly inherit them unchanged.

The Operational Details That Decide Outages

Four mechanics separate agent deployments that survive incidents from the ones that create them:

  • Graceful shutdown is a state machine, not a flag. On termination, the worker stops starting new steps, finishes or checkpoints the current one, persists state, acknowledges the message, and exits inside the drain window. Anything less turns every deploy into a small data-loss event. Termination grace periods must be sized to the longest step, not the average one, or the platform will kill runs mid-step that your code believed were draining.
  • Autoscaling follows two signals, not one. Queue depth tells you backlog; oldest-message age tells you urgency. Depth alone overreacts to bursts; age alone underreacts until users have already waited. Set the per-worker concurrency cap from provider limits: with a rate ceiling of R requests per minute and a worst case of C calls per run-minute, each worker holds at most R divided by C concurrent runs. Then cap the pool so the maximum fleet spend is a number the business approved in advance.
  • Agent versions deploy blue-green, evaluated. Prompts, tool contracts, and models change behavior, so ship them like releases: run the new version beside the old, shift a slice of traffic, compare success rate and cost per successful task on the dashboards, then widen or roll back. The evaluation suite is the gate; the deploy pipeline is its enforcement arm.
  • Status is a contract from second zero. Every run gets a stable id and a status endpoint with explicit states, idempotent reads, completion events for subscribers, and a dead letter path that keeps the trace attached to runs that exhaust their attempts. Users forgive waiting they can see. They do not forgive work that vanishes.

State, Queues, and Resumability

Run state belongs in a durable store, never in worker memory. The pattern that works is unglamorous and proven:

  • Checkpoint per step. After every loop iteration, persist working state. A crashed worker resumes from the last checkpoint instead of restarting the task and repeating side effects the user already suffered.
  • Idempotent writes. Every tool with side effects accepts an idempotency key derived from the run. Queue redelivery and worker crashes become non-events instead of double refunds. The full model is agent state persistence.
  • A status endpoint from second zero. Runs move through explicit states: queued, running, waiting for approval, completed, failed. Clients poll or subscribe, and the user always knows what happened to their request.
  • Backpressure and a dead letter path. Queues absorb bursts; when the queue itself fills, shed and alert rather than scale spend without limit. Stuck runs land in a dead letter queue with their trace attached, where a human can decide instead of a retry loop can guess.

Security Boundaries in Production

The deployment inherits every security principle from tool design, applied at infrastructure level:

  • Secrets live in a manager, not in the loop. The orchestrator and workers hold short-lived, scoped credentials. An agent process with admin credentials is an over-permissioned tool at operating-system scale.
  • Egress is controlled. Workers call exactly three things: the model provider, the tool gateway, and the retrieval layer. Anything else, the sandbox denies. What that restriction actually means in practice is tool sandboxing and permission models.
  • Tenant isolation is enforced by the stores. State, retrieval, and telemetry all scope by tenant at the data layer, so no prompt can ask its way across the boundary.
  • The tool gateway is the audit surface. Every action, every authorization decision, every approval, logged once, centrally, through the choke point the architecture exists to create.

How Real Systems Do This

  • Agents embed into existing platforms, not beside them. My content localization agent runs inside the CMS: the platform is the workflow, the model-driven steps are the flexible parts, and the agent capability ships where content already lives instead of beside it.
  • Hybrid API everywhere. Accept, validate, return a run id, execute asynchronously. Even fast agents ship this way, because the pattern costs little and makes long runs possible later.
  • Workers are stateless; state is a service. Deploys drain workers at step boundaries, and the queue redelivers. Lost work on deploy is treated as a bug, not a tradition.
  • The tool gateway is a shared platform service. Authorization, rate limits, and audit live in one place, serving every agent in the organization.
  • Cost dashboards run per agent version, with alerts on cost per successful task, so a prompt change that doubles spend shows up in hours, not invoices.
  • Agent changes ship like releases. Prompt, model, and tool-contract changes roll out progressively, with evaluation gates and rollback, because behavior changes are behavior changes regardless of which file they live in.

Decision Framework

Answer these before the first deploy, in order:

  1. How long can a run take? Seconds, request-scoped is defensible. Anything longer becomes a job with checkpoints, a status endpoint, and idempotency.
  2. What happens if the process dies mid-run? Checkpoint cadence and resume must have an answer. “It restarts” is the wrong answer when tools have side effects.
  3. What are the per-user and per-tenant cost ceilings? Enforce at the gateway, before spend happens.
  4. How many concurrent runs per worker? Derive from provider rate limits and budget enforcement, not from CPU. Then cap the worker pool and make overflow queue, alert, and shed.
  5. Where do retries stop being safe? Every write tool gets an idempotency key; every retry policy respects it. The full reliability ladder is reliability engineering for agents.
  6. What can leave the process? Credentials scoped, egress restricted, telemetry metadata-only by default.
  7. How does a deploy behave mid-run? Drain at step boundaries, checkpoint, redeliver. If this is not written down, every release is a small outage.

When NOT to Use This

  • Single-user internal tools. A service, a cron job, and a database can be the right architecture for a while. The full topology earns its cost at multi-user, multi-tenant scale.
  • Prototypes still changing weekly. Deploy minimal: one service, tracing, a step budget. Elaborate infrastructure around an unstable loop is scaffolding on sand.
  • Strictly interactive, sub-second products. Agent runs do not belong on that path at all. Rethink the feature rather than deploying it into a latency budget it will violate.
  • No defined cost ceilings. If the business cannot name a per-user spend limit, the platform is not ready to face users. The first invoice will name one for you.

Common Mistakes

  • Long runs inside requests. The cost: timeouts that kill runs mid-investigation, client retries that restart everything, and duplicate side effects from the retry.
  • API-server concurrency settings on agent workers. What you get: token-bound runs strangling workers built for CPU-bound traffic, and queues nobody noticed filling.
  • Idempotency deferred. Where it lands: the first queue redelivery or worker crash becomes a duplicated refund, discovered by its recipient.
  • Admin credentials in the loop process. The consequence: every prompt-injection discussion becomes an incident-response discussion.
  • Deploys without drain discipline. What follows: every release quietly kills in-flight runs, and users learn not to trust long tasks.
  • No per-user cost caps. The price: denial of wallet, delivered by your most enthusiastic customer.

Key Takeaways

  • A production agent is a small distributed system: gateway, orchestrator, tool gateway, state store with checkpoints, queue and workers, observability.
  • Synchronous for runs measured in seconds. Asynchronous jobs with checkpoints, status endpoints, and idempotency for everything longer.
  • The gateway is a cost control: per-user caps, run limits, and validation before any model spend.
  • Workers are token-bound: low concurrency per worker, queue-depth scaling, cost-aware autoscaling policies, graceful drain at step boundaries.
  • Security is structural: scoped credentials, controlled egress, tenant-isolated stores, and the tool gateway as the single audit surface.
  • Ship behavior changes as releases: prompts, models, and tool contracts all gate on evaluation, with rollback.

FAQ

How do you deploy an AI agent in production?

As a small distributed system: an API gateway for auth, validation, rate limits, and cost caps; an orchestrator service running the loop with budgets; a tool gateway for authorization and audit; a state store with per-step checkpoints; a queue and workers for long runs; and tracing throughout. Fast runs can stay request-scoped, and everything longer becomes a job.

Should an AI agent run synchronously or asynchronously?

Runs that finish in a few seconds can live inside a request. Anything longer belongs in a background job with checkpointed state, a status endpoint, and idempotent writes, because request timeouts will otherwise kill runs mid-task and retries will duplicate side effects.

How do you scale AI agents?

Scale workers on queue depth, with concurrency caps per worker derived from provider rate limits, because agent runs are token-bound rather than CPU-bound. Cap the worker pool so scaling cannot multiply spend without limit, and drain workers at step boundaries so deploys do not destroy in-flight runs.

What infrastructure do AI agents need?

A gateway, an orchestrator service, a tool gateway, a durable state store, a queue with workers, retrieval if the task needs evidence, and an observability backend. Plus the unglamorous essentials: a secrets manager, scoped credentials, egress controls, and a dead letter path for stuck runs.

How do you handle long-running agent tasks?

Checkpoint state after every step, expose a status endpoint with explicit run states, give every write tool an idempotency key, resume from the last checkpoint after crashes, and let approvals pause runs as persisted state waiting for input rather than blocked threads. Users tolerate waiting they can see.

Conclusion

Agent deployment is service deployment with three additions the agent brings itself: metered runs that need cost ceilings, nondeterministic decisions that need traces, and real actions that need permission gates. None of the additions is novel. All of them are old disciplines wearing new labels, which is either discouraging or a relief, depending on how much distributed-systems experience the team has.

The architectures that work are boring on purpose: gateways that cap, workers that checkpoint, queues that buffer, and one choke point where every action is authorized and recorded. The interesting behavior lives inside the loop. Everything around it should be forgettable.

Deploy the agent like any distributed system, cap it like a metered utility, and gate it like it can spend money. Because it can.

Last updated on 6 October 2026

Share this article

Leave a Reply

Your email address will not be published. Required fields are marked *