Tool Sandboxing and Permission Models: Least Privilege for Agents
Tool sandboxing and permissions for AI agents: least privilege, role tiers, capability isolation, and sandboxes that actually limit what tools can reach.
“The agent needed access” is how most agent security stories begin. An integration goes faster with broad credentials, a tool gets an admin token because scoping it took an afternoon, and a system designed to be persuaded by text now holds keys sized for a person it is not.
Permission models for agents follow one principle with unusual clarity: least privilege, because the caller is nondeterministic. A human with broad permissions at least behaves consistently when tricked once. An agent can be tricked differently on every run, at machine speed, through channels humans do not read.
This article covers the permission models that work for agents, how to design tiers, where enforcement lives, what sandboxes actually restrict, and how blast radius gets reduced until an agent’s worst day is a bad afternoon instead of an incident.
The Principle: Least Privilege for a Nondeterministic Caller
Least privilege is old doctrine: every subject gets the minimum access required for its function. Agents make it non-optional for three reasons:
- The caller can be persuaded. Prompt injection, covered in its own article, means instructions can arrive through data. Permissions are the layer that does not care whether the model was convinced.
- Decisions are frequent and varied. A human admin makes hundreds of privileged decisions a year; an agent can make thousands a day, each one a fresh chance to be wrong in a new way.
- Actions are the product. The permission model is not a compliance checkbox around an agent; it is a core component of the agent’s design, as load-bearing as the tool contracts in tool use and function calling.
OWASP names the failure mode directly: excessive agency (granting LLM-based systems more permissions, autonomy, or budget than their function requires) is a top-ten risk in its own right: OWASP Top 10 for LLM Applications.
Permission Models, Compared
| Model | How it decides access | Fits agents when | Weakness |
|---|---|---|---|
| Role-based | Role holds permissions; caller holds role | Coarse trust tiers: reader, writer, approver | Roles drift toward “admin” over time |
| Scope-based | Per-resource, per-action scopes on tokens | Fine-grained control on named resources | Scope sprawl; needs disciplined design |
| ACL-based | Per-object lists of permitted subjects | Document and record level access | ACL maintenance at scale |
| Capability-based | Unforgeable references to specific rights | Delegation: handing a run exactly one capability | Less familiar tooling in most stacks |
Production agents usually combine the first two: role tiers for the trust ladder, scopes for the resources within each tier. The capability idea still earns a place conceptually: the ideal grant to a run is a short-lived, narrow token that names one capability and expires with the run, which is also the answer to “do agents need their own credentials”: yes, scoped ones, never shared human credentials.
Designing Permission Tiers
The trust ladder for agents has three rungs, and every tool belongs on exactly one:
- Read-only. Lookups, searches, status checks. Default for every new agent, with full audit and no gate.
- Scoped writes. Records changed, messages sent, files written. Role-gated, quantity-capped, idempotency-keyed, and often approval-gated on top.
- Irreversible actions. Payments, deletions, external legal or infra effects. Named human approver, no exceptions, per the approval design.
Tier 1 read-only tools default for every agent
|
+-- elevate only on need, with review
v
Tier 2 scoped writes role-gated, capped, keyed
|
+-- never without a named human
v
Tier 3 irreversible actions approver, audit, no exceptions
Two implementation rules keep the ladder honest. First, roles attach to runs, not to processes: a support run holds the support role, and elevation is an explicit workflow event with an audit entry, never a config drift. Second, credentials are per-run, short-lived, and scoped: the run gets a token that names its capabilities and expires when it ends. A token that outlives its run is a permission you no longer control.
The Tool Gateway as Enforcement Point
Enforcement lives in one place, or it lives nowhere. The tool gateway, positioned between the orchestrator and every external system in the deployment architecture, is where four functions meet:
- Authorization: role and scope checks, independent of what the model intended.
- Validation: contract checks on arguments before execution, the enforcement twin of the tool contracts.
- Rate and quantity limits: per caller, per tool, per time window.
- Audit: every call logged with caller, role, arguments hash, decision, and outcome.
The gateway is the piece that makes prompt injection survivable rather than fatal: the attacker may convince the model, but the gateway answers a different question, “may this caller do this,” and it does not read prose.
Sandboxing: What It Actually Restricts
Sandboxing gets described in marketing language more often than in boundary language, so state the specification plainly. A container-based sandbox for an agent’s code or file tools restricts reach along four axes:
- Filesystem: mounted volumes only. The sandbox sees its scratch directory and the inputs you mount, and nothing else.
- Process: restricted process spawning, with syscall profiles limiting what executed code can ask the kernel for.
- Network: egress allowlisting. No allowlist, no egress. This single control closes most exfiltration routes and much of the SSRF attack surface.
- Resources: CPU, memory, and time limits, so a compromised step cannot exhaust the host.
Equally important, the sandbox does not restrict: the model’s behavior, its willingness to follow smuggled instructions, or anything about the secrets you mount inside it. A sandbox is a reach boundary, not a behavior boundary, and confusing the two is how teams end up with containers that have everything mounted and a diagram that says “sandboxed.” The container primitives that make these boundaries real are documented thoroughly in the Kubernetes documentation, and agent workloads inherit them unchanged.
Blast Radius Reduction in Practice
The design exercise is one question, asked per tool: given the worst compromised decision this tool enables, what is the maximum damage? Then reduce the answer until it is acceptable:
- Quantity caps. Maximum rows per query, maximum refund amount, maximum recipients per send.
- Destination allowlists. Email domains, write paths, API endpoints. Novel destinations fail closed.
- Idempotency keys. One per run and step, so retries and redeliveries cannot multiply the effect.
- Time-boxed credentials. Capabilities that expire with the run limit how long a leaked token matters.
- Per-tenant scoping. Every credential and every query bound to one tenant, enforced in the store, not the prompt.
Specifying a Sandbox Like an Engineer
A sandbox specification is a list of denials, not a product name. The honest spec sheet for an agent code or file tool:
- Non-privileged execution. The workload never runs as root inside or outside the container; privilege escalation is a bug class you exclude on purpose.
- Syscall allowlisting. Kernel-surface profiles restrict what executed code may ask the operating system for; execution tools need a short list of syscalls, not all of them.
- Read-only mounts for inputs, ephemeral scratch for outputs. The sandbox writes only to a scratch directory that dies with the run.
- No secrets in the environment by default. Anything mounted is reachable by definition; the base image carries no credentials, and specific tools get specific scoped tokens.
- Time and resource ceilings. CPU, memory, and wall-clock caps, so a compromised step dies of exhaustion instead of succeeding.
Each line is testable: attempt the denied thing in CI, and the test fails if the sandbox permits it. That is what separates a boundary from a diagram.
Designing Scopes That Survive Review
The same honesty applies to scopes. The MCP authorization guidance, written for the protocol but generalizing cleanly to any tool layer, recommends: start from a minimal baseline of low-risk discovery and read operations; elevate per operation with step-up challenges when a privileged call is first attempted; never publish omnibus scopes like “all” or “full-access”; version scope meanings so silent changes cannot re-grant old tokens; and enforce server-side, because a token’s claims are an input, not an authorization decision. The guidance lives with the protocol’s security documentation: MCP Security Best Practices.
A workable naming convention pairs resource with action per tool, “orders:read” versus “orders:refund”, so the audit log reads like a sentence and reviewers can spot an over-broad grant without a meeting. Then run the compromised-tool thought experiment per scope: if a hostile instruction convinced the model to call this tool with plausible arguments, what is the worst outcome? For an email tool with an allowlisted recipient domain, it is one embarrassing message. For the same tool with unrestricted sending, it is a data breach with a friendly subject line. The difference is one scope decision made before launch, not after.
How Real Systems Do This
- New agents ship read-only. Write tiers arrive after run history justifies them, and the upgrade is a reviewed permission change, not a quiet config edit.
- Per-run scoped tokens come from a secrets manager. No agent process holds a human’s credentials, and no credential outlives its run.
- Code execution runs egress-off by default. The sandbox gets a scratch filesystem, mounted inputs, and an empty network allowlist until a requirement names a destination.
- Permissions are re-reviewed on a schedule. Role drift is checked quarterly the way access reviews work everywhere else, because agents accrue permissions exactly like services do: quietly.
- Injection tests validate the walls. Red-team scenarios from prompt injection defense are run against the permission model, because a boundary that has never been tested is a hypothesis.
Decision Framework
For each tool, in order:
- What does it touch? Name the resources, systems, and data. This inventory is the security surface, and it cannot be delegated to the model.
- Which rung does it sit on? Read, scoped write, or irreversible. If the answer is uncomfortable, the tool scope is wrong, not the ladder.
- What credential does it use? Per-run, scoped, short-lived. A shared long-lived token is a finding, not a configuration.
- What does its sandbox restrict? Filesystem, process, network, resources, written down. If nobody can answer, the sandbox is decorative.
- What are the caps and allowlists? Quantities, destinations, and rates, all failing closed.
- Which actions require a human? Irreversible ones, by definition, and any write whose worst case exceeds the tolerance.
- Can any action be reconstructed from the audit log? Caller, role, arguments, decision, approver, outcome. If the incident review would guess, the audit is incomplete.
When NOT to Use This
- Read-only agents over public data. Hygiene applies, least privilege and audit, but the full tier and sandbox machinery is for agents that hold real permissions.
- Pre-production prototypes. Start read-only from day one, because it is nearly free, and defer the rest until the system earns production.
- Target systems without fine-grained scopes. If the downstream API offers only an admin token, the choice is to build the scoping layer or not expose the tool. “Broad token plus careful prompting” is the option that does not exist.
Common Mistakes
- Shared admin credentials, temporarily. The price: a permanent blast radius attached to a system designed to be persuadable.
- Sandbox theater. A container with everything mounted and unrestricted egress. In practice: a boundary that exists in diagrams and nowhere else.
- Roles that only escalate. The result: every agent drifts toward admin over a quarter, because no permission request was ever denied.
- Long-lived credentials. The damage: a leaked token that outlives the incident that leaked it.
- Audit without correlation. The fallout: logs that cannot answer “which run did this,” and incident reviews that end in attribution arguments.
- Behavior controls sold as reach controls. The cost: prompt wording standing where a permission check belongs, which is the exact mistake the injection literature keeps documenting.
Key Takeaways
- Least privilege is non-optional when the caller is nondeterministic and persuadable. Permissions are the layer that does not read prose.
- Every tool sits on a rung: read-only, scoped write with role and caps, or irreversible with human approval.
- Enforce at a tool gateway: authorization, validation, limits, and audit in one choke point, independent of model intent.
- Credentials are per-run, scoped, and short-lived. Agents never hold shared human credentials.
- A sandbox restricts reach (filesystem, process, network, resources), not behavior. Specify what it actually restricts, or admit it is decoration.
- Blast radius is reduced with caps, allowlists, idempotency keys, and tenant scoping, all failing closed.
- Audit everything through the gateway, correlated by run, and re-review permissions on a schedule before drift becomes architecture.
FAQ
What is tool sandboxing for AI agents?
Running an agent’s executed tools, especially code execution and file operations, inside a constrained environment: a container with mounted-only filesystem access, restricted process spawning, allowlisted network egress, and resource limits. It bounds the reach of any single tool call, including one resulting from a successful prompt injection.
What is least privilege for AI agents?
Giving each tool and each run the minimum access its function requires, on the assumption that the caller can be persuaded. Reads by default, writes behind roles and caps, irreversible actions behind human approval, and credentials that are scoped, per-run, and short-lived.
How do you restrict what an AI agent can access?
Structurally, at the tool gateway: role tiers per tool, scoped per-resource permissions, quantity caps and destination allowlists, tenant filters enforced in the stores, and per-run tokens. Prompt instructions restrict nothing; boundaries enforced in code restrict everything that matters.
What does a sandbox actually restrict?
Reach, not behavior. Filesystem access to mounted volumes, process spawning through syscall restrictions, network through egress allowlists, and host resources through limits. It does not restrict the model’s willingness to follow instructions, and it restricts nothing you mount inside it.
Do AI agents need their own credentials?
Yes, and only their own. Per-run, scoped, short-lived tokens from a secrets manager, sized to the task. An agent running on a human’s credentials or a shared admin token is not a convenience; it is a standing incident with a user interface.
Conclusion
Agent security is permission engineering. The model will sometimes be tricked, the tools will sometimes fail, and the runs will sometimes surprise you. The systems that survive are the ones where surprise has a ceiling: every capability scoped, every write keyed and capped, every irreversible action approved by a human, and every action reconstructable from the audit log.
The work is unglamorous and mostly familiar, which is the good news: access reviews, scoped tokens, allowlists, and audit trails are decades-old disciplines. The agent era did not invent them. It removed the option of ignoring them, because it built a system whose judgment is a service that can be called by anyone who controls its inputs.
An agent with unrestricted tools is not autonomous. It is underconstrained.
Last updated on 6 October 2026
