RAG for Agents: Retrieval That Adapts to the Task
RAG for agents vs chatbot retrieval: task-aware queries, grounding, access control, freshness, and retrieval as a tool the agent decides to call.
Ask a chatbot a question and retrieval happens once: fetch three passages, answer. Ask an agent to investigate a billing discrepancy and retrieval happens five times, shaped differently each time, because each step needs different evidence.
RAG in an agent is a capability the loop calls, not a pipeline wrapped around the model. That inversion changes the design. Queries are generated from task state. Results must carry provenance and respect access control. And retrieval quality becomes agent quality, because the evidence the model sees decides the actions the model takes.
If you have not built a retrieval pipeline before, start with how to build a RAG app in Python, which this article assumes. This article covers retrieval as an agent capability: the pipeline, task-adaptive queries, grounding, permissions, freshness, and the latency and cost you are signing up for. It builds on the memory model in agent memory: retrieval is infrastructure that feeds the loop, distinct from what the system remembers about users and runs.
RAG in an Agent Is Not Chatbot Retrieval
The original retrieval-augmented generation paper (Lewis et al., 2020) framed the idea precisely: combine a model’s parametric memory with a non-parametric store (a dense index accessed by a neural retriever) for knowledge-intensive generation. That framing, and the observation that models cannot precisely update or attribute their internal knowledge, is why retrieval exists at all: Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks.
What changed with agents is where retrieval sits. In a chatbot, retrieval runs once per turn, feeding the answer. In an agent, retrieval is a tool inside the loop, called zero or more times per run, feeding the next decision. The distinction is architectural, and it shows up in six places:
| Dimension | Chatbot RAG | Agent RAG |
|---|---|---|
| When retrieval runs | Once per turn | Whenever the loop decides |
| Query source | The user message | Task state and current subtask |
| What results feed | The answer text | The next decision or tool call |
| Access control | Session-level | Per-call tenant and permission filters |
| Failure impact | A wrong passage | A wrong action chain built on wrong evidence |
| Cost profile | Predictable | Scales with iterations |
One more agent-specific property: retrieved content is untrusted input. A document in your index can contain instructions, and an agent that acts on retrieved text is a target for indirect prompt injection, a risk class covered in prompt injection in agents. Retrieval design and injection defense are the same project.
The Retrieval Pipeline for Agents
Whether the agent calls it once or five times, each retrieval pass is the same five-stage pipeline:
- Query generation. Turn the current task state into one or more retrieval queries. This is a generation step, not a copy step.
- Retrieval. Execute the queries against the index or search backend, usually semantic search, keyword search, or a hybrid.
- Ranking and filtering. Score candidates for relevance, apply permission and tenant filters, and drop what does not belong in this run.
- Context assembly. Project the winners into bounded, well-labeled passages with source and timestamp, sized to a context budget.
- Model decision or response. The evidence enters the loop: it grounds the next tool call, the final answer, or a follow-up retrieval.
The stages are separable, testable, and independently improvable. When retrieval quality drops in production, the failure lives in exactly one stage, and the trace should say which: a bad query, a bad index, a bad filter, or a bad assembly.
Task-Adaptive Query Generation
The most common quality gap in agent RAG is queries that echo the user instead of serving the step. The user wrote “why is my bill different this month.” The step that needs retrieval is “compare the current plan’s overage rates with the customer’s usage pattern,” which is a different query, and neither is the user’s sentence.
Practical techniques, in increasing cost order:
- Step-derived queries. Generate the query from the current subtask and state, not from raw user text. The orchestrator can also derive queries deterministically from entities it already extracted.
- Reformulations. Generate two or three variants per step to survive vocabulary mismatch, and merge the results.
- Hybrid search. Combine semantic search with keyword matching, since identifiers, error codes, and SKUs do not embed well.
- Multi-hop. Let step one’s findings generate step two’s query. This is where retrieval-as-tool shines, and where a fixed pipeline cannot follow.
Grounding and Citations
An agent that acts needs evidence discipline a chatbot never required. Every claim in a final answer should trace to a retrieved source, and the observation layer must carry that provenance: source document, section, and retrieval time.
Two enforcement points make grounding real. First, citations in the output: the answer names its sources, and the evaluation suite checks that named sources actually support the claims, which is a measurable property covered in agent evaluation. Second, provenance in the trajectory: when an action was taken “based on policy,” the trace shows which passage, from which index, retrieved when. An unsourced claim from an agent that can act is a rumor with tool access.
The product solution architect agent I architected and delivered does the same class of work: it synthesizes business intelligence from public and enterprise sources for consultative recommendations, and in that domain a recommendation without traceable provenance is a guess wearing a suit. Grounding is what makes such a product defensible in front of a client.
Access Control and Multi-Tenancy
Retrieval that ignores permissions is a data leak with a search box. The enforcement belongs in the retrieval layer, applied as query-time filters, not in the prompt:
- Tenant isolation. Every query carries the caller’s tenant, and the index enforces it, so no prompt can ask its way into another tenant’s documents.
- Document-level ACLs. Role and user entitlements become metadata filters. A passage the user could not open in the source system must not surface in a retrieval result.
- Query logging. Every retrieval, with filters and results, is an audit event. The question “how did the agent know that” must have an answer.
This is also the injection-adjacent boundary: retrieved documents are content, not instructions, and the defenses against content-borne attacks are architectural. OWASP’s LLM risk list treats both sensitive disclosure and prompt injection as top-tier concerns for exactly this reason: OWASP Top 10 for LLM Applications.
Freshness, Chunking, Latency, and Cost
Three operational knobs decide whether retrieval helps or quietly degrades the agent.
- Freshness. Every retrieved passage should carry a timestamp, and the agent should see it. Index lag is fine for policy documents and fatal for prices, statuses, and inventory. Fast-changing facts may belong to a live lookup tool instead of the index, and retrieval observations should make staleness visible, not hidden.
- Chunking. Small chunks give precision but fragment context; large chunks give coverage but bury signal in noise. Chunk with metadata (source, section, date, entitlements) so filters and citations work. The assembly stage should project passages to what the step needs rather than dumping full chunks.
- Latency and cost. Retrieval is a tool call: it adds a hop per iteration, and multi-hop runs multiply it. Budget retrieval calls like model calls, cache where the corpus is stable and the query repeats, and treat the context that results enter as a priced resource, a discipline covered in context window management.
A Compact Retrieval Tool
The example below is illustrative: wire INDEX to your search backend. The structural points are the ones to keep: permission filters applied inside the tool, bounded snippets, provenance and timestamps in every result.
import json
class SearchIndex:
"""Stub search backend. Replace with your vector or keyword index."""
def search(self, query, top_k, tenant, allowed_sources):
return [] # a real backend returns document objects
INDEX = SearchIndex()
def search_knowledge(args: dict) -> str:
"""Retrieval tool body over the stub index above."""
query = args["query"]
docs = INDEX.search(
query,
top_k=args.get("top_k", 5),
tenant=args["tenant"], # enforced, not requested
allowed_sources=args.get("allowed_sources", []),
)
if not docs:
return json.dumps({"results": [], "note": "no documents matched"})
return json.dumps({"results": [
{
"title": doc.title,
"snippet": doc.text[:500], # bound the observation
"source": doc.source, # provenance for citations
"updated_at": doc.updated_at, # freshness signal
}
for doc in docs
]})
Notice what the tool refuses to do: it does not trust the caller to filter results, it does not return unbounded text, and it does not strip the metadata the agent needs to cite and date its claims. Retrieval tools earn their place by being the only component that touches the index, which makes the access rules one place instead of many.
How Real Systems Do This
- Support agents retrieve policy and account state with ACLs, and their replies to users include citations, because “per policy” without a source is a complaint waiting to happen.
- Research agents run multi-hop retrieval, each hop generated from the previous findings, with citations required in the final artifact. The retrieval trace is the audit trail.
- Enterprise deployments put a retrieval gateway in front of the index, so tenant filters, rate limits, and query logging happen in one place for every agent, every run.
- Fast-changing facts go to live tools, not the index: prices, stock, and order statuses come from API calls, while the index serves the stable knowledge. The agent sees the difference in the observation labels.
- Mature teams evaluate retrieval directly: test tasks with known expected evidence, and a metric for “right evidence surfaced at the right step.” Silent corpus drift is caught by the suite before users catch it.
Decision Framework
Answer these before wiring a retrieval tool into an agent:
- Does the task need external, changing knowledge? If the facts are stable and fit in the prompt, include them and skip the index.
- Is one retrieval per run enough? Then a fixed pipeline before the loop is simpler and cheaper than retrieval-as-tool.
- Do queries need to adapt mid-run? Multi-hop investigation is the case that justifies retrieval-as-tool.
- Are permissions enforced in the index? If not, stop. Retrieval before access control is a leak, not a feature.
- Does the output need citations? Then provenance must ride along in every observation, and grounding must be in the eval suite.
- What freshness does the task require? Index lag acceptable, or does the agent need live lookups for the fast-changing parts?
- Is retrieval budgeted? Hops per run, tokens per result, and context share should all have numbers before launch.
When NOT to Use This
- Stable knowledge that fits in the prompt. A few paragraphs of policy in the system context beat an index round trip and its failure modes.
- A low-quality corpus. Retrieval amplifies whatever it finds. If the source material is wrong or stale, the agent will cite it with confidence, which is worse than not citing.
- Latency-critical paths. Retrieval adds a hop per iteration. Hot paths need cache hits or precomputed context.
- Undefined access model. If nobody can say who may read which document, do not connect a model to the corpus. The leak will be fluent.
Common Mistakes
- Retrieval as an answer engine. In practice: stuffed context, no citations, and a model that interpolates between passages it never verified.
- No access control in the index. The result: cross-tenant leakage, discovered by the first curious user or the first auditor.
- Oversized passages. The damage: context bloat on every retrieval, and the signal the step needed buried in the noise the step did not.
- Ignoring freshness. The fallout: the agent citing a price, a plan, or a status that expired last quarter, with total confidence.
- One fixed query per run. The cost: multi-hop questions answered from single-hop evidence, with the gaps filled by the model’s imagination.
- Unmeasured retrieval. What you get: quality decaying silently as the corpus drifts, and the first measurement arriving as a user complaint.
Key Takeaways
- In an agent, retrieval is a tool inside the loop, not a pipeline around the model. Queries adapt per step; results feed decisions.
- The five stages are separable and testable: query generation, retrieval, ranking and filtering, assembly, decision. Trace failures to a stage.
- Grounding requires provenance in observations and citations in outputs. Unsourceable claims are unauditable claims.
- Access control lives in the retrieval layer as query-time filters. Prompts are not permission systems.
- Freshness is visible: timestamps in observations, live tools for fast-changing facts, index lag acknowledged rather than hidden.
- Budget retrieval like any tool call: hops, tokens, and context share. Measure it with expected-evidence test sets.
FAQ
What is RAG for agents?
Retrieval-augmented generation deployed as an agent capability: a search tool the loop calls when a step needs external evidence, with queries generated from task state, results carrying provenance and permissions, and citations required in outputs. It feeds the agent’s decisions, not just its answers.
How is agent RAG different from chatbot RAG?
A chatbot retrieves once per turn from the user’s message and answers. An agent retrieves when it decides to, from queries derived from its current subtask, and the results feed the next tool call or a follow-up retrieval. Access control and cost profiles are correspondingly stricter.
Should an agent always retrieve before answering?
No. Retrieval is a decision like any tool call. If the task’s knowledge is stable and small enough for the prompt, retrieving adds latency, cost, and injection surface for nothing. Agents should retrieve when the evidence they need is external, changing, or too large to include.
How do you stop retrieved content from hijacking an agent?
Treat retrieved text as untrusted data, never as instructions: keep tool results and instructions structurally separated, label provenance, validate outputs against policy, and let permission boundaries, not prompt wording, constrain what retrieved content can achieve. The full defense-in-depth treatment is in prompt injection for agents.
How do you evaluate retrieval quality in an agent?
Build test tasks with known expected evidence and measure whether the right documents surface at the right step, then check grounding: whether cited sources actually support the claims made. Track per-stage failures so quality decay attributes to query generation, the index, filters, or assembly.
Conclusion
Retrieval is the agent’s evidence system. Chatbot RAG could treat evidence as garnish for an answer; an agent cannot, because the same passages justify actions, and actions have consequences, audit trails, and incident reviews.
The engineering that makes it safe is unglamorous and learnable: adaptive queries, bounded assembly, provenance everywhere, permissions in the index, freshness made visible, and a test suite that notices when the corpus drifts. None of it is optional in production, and all of it compounds.
Retrieve what the step needs, prove where it came from, and never trust it more than its source.
Last updated on 4 October 2026
