Tracing Agent Runs: LangSmith, Langfuse, and OpenTelemetry

Tracing agent runs: spans for model calls, tools, retrievals, and state changes in LangSmith, Langfuse, or OpenTelemetry, without leaking secrets.

Executive Summary: Trace the whole agent run, not only model calls: one root span with nested spans for steps, tool calls, retrievals, retries, and failures, each with timing, status, and token usage. This post explains why prompt and tool content should be exported only when redacted, and how to choose between LangSmith, Langfuse, and plain OpenTelemetry for your stack.

Logs answer “what lines ran.” Traces answer “what happened.” For an agent, the second question is the only useful one: a failed run is a story of decisions, tool calls, observations, and state changes, and the trace is the only artifact that tells it in order.

An agent run is a distributed workflow that fits inside one process. It has a root span, nested spans (for steps, model calls, tool executions, retrievals, retries, and state transitions), and a completion. Distributed-tracing tools have modeled exactly this shape for years, and the agent ecosystem has converged on it.

This article covers tracing agent runs: the span taxonomy, what to record and what never to record, how LangSmith, Langfuse, and OpenTelemetry compare, and how traces become the debugging and evaluation backbone of production agents.

An Agent Run Is a Distributed Workflow

The mental model comes from distributed tracing, and it fits agents exactly. A run begins with a root span for the whole task, and everything that happens becomes a child span with timing, status, and attributes. The hierarchy for a typical agent run:

root run: resolve support ticket #4821
|
+-- agent step 1
|   +-- model call (provider, model, tokens in/out)
|   +-- tool call: lookup_order (args hash, latency, status)
|   +-- state transition: working_state updated
|
+-- agent step 2
|   +-- retrieval: query, filters, top-k, latency
|   +-- model call
|   +-- tool call: create_refund (approval: human-42)
|   +-- retry: create_refund (attempt 2, transient timeout)
|
+-- completion: final answer, total cost, 14.2s

Every span answers the same questions: what happened, how long did it take, did it succeed, and what did it cost. The root answers the only question leadership asks: did the task succeed, and what did it cost.

The Span Taxonomy for Agents

Nine span types cover an agent run completely. Each answers one operational question:

Span What it records The question it answers
Root run Task, session id, final outcome, totals Did the task succeed, and what did it cost?
Agent step Iteration boundary, context size at entry Where did the run spend its steps?
Model call Provider, model, tokens, latency Which decisions used which model, at what price?
Tool call Tool name, argument hash, status, latency Which actions ran, and how did they behave?
Retrieval Query hash, filters, top-k, latency What evidence entered the run?
State transition What changed in working state, schema version What did the run believe, and when?
Retry Attempt number, trigger, backoff Where does flakiness live?
Failure Error type, failing span, stage What breaks most, and where?
Completion Final outcome code, total cost, duration The bottom line, per run

The taxonomy is deliberately vendor-neutral. LangSmith, Langfuse, and OpenTelemetry all express it, with different names and different conveniences, which is the strongest evidence that the shape is real.

What to Record on Every Span

Four attribute groups, all cheap and all load-bearing:

  • Identity. Run id, session or conversation id, agent identity and version, model and provider, and a tenant or user identifier, hashed where appropriate. The version attribute is the one teams skip and then miss: without it, no regression can be attributed to the change that caused it.
  • Timing and status. Start, end, duration, and a status that distinguishes success, failure, timeout, and cancellation.
  • Economics. Token usage per model call and estimated cost per span. Cost incidents start as anomalies in these fields.
  • Outcome codes. Structured results: tool status codes, approval decisions, escalation events, final outcome. Codes, not prose, so dashboards and alerts can aggregate them.

The Privacy Problem

Here is the uncomfortable truth about agent tracing: the most useful data to record (prompts, tool arguments, tool results, retrieved documents) is also the most dangerous. Customer content, credentials pasted into arguments, personal data in retrieved records. A telemetry pipeline that exports all of it by default has quietly built a second, less protected copy of your most sensitive data, and shipped it to a third-party backend.

The controls that make tracing safe are not optional accessories:

  • Content off by default. Record metadata, hashes, counts, and codes on every span. Capture content only for specific span types, behind explicit opt-in.
  • Redaction at emission. Scrub secrets and personal data before export, not after storage. The emission point is the only place with the context to redact correctly.
  • Retention limits. Telemetry stores accumulate the full history of what your agents saw. Retention policy is a compliance control, not a cost knob.
  • Access control on the store. The tracing backend becomes one of your most sensitive databases. Treat its access model accordingly.

The discipline is worth stating as a rule: trace everything the run does, record almost nothing it says.

LangSmith, Langfuse, and OpenTelemetry

All three options implement the same taxonomy; they differ in ownership and integration.

Consideration LangSmith Langfuse OpenTelemetry
Positioning Managed observability and evaluation platform Open-source AI engineering platform Vendor-neutral instrumentation standard
Integration Deep LangChain and LangGraph pairing, many framework integrations 100+ integrations, sessions, agent graph visualization Instrument once, export to any backend
Standards Platform-specific trace model Built on OpenTelemetry, reducing lock-in Defines GenAI semantic conventions, including agent and MCP spans
Evaluation Tracing, datasets, judges, monitoring in one Tracing, prompt management, evals, annotation queues None: bring your own evaluation layer
Operations Vendor-run, priced by usage Self-hostable or cloud You own the pipeline and backend
Choose when LangChain-heavy stack, managed workflow, eval-first culture Self-hosting, compliance control, OTel alignment Multi-vendor estates and long-term portability

The official documentation for each is the ground truth for capabilities, and all three move quickly: LangSmith docs, Langfuse docs, and the OpenTelemetry documentation. One structural note from the OTel side: its semantic conventions for generative AI now cover agent spans and MCP spans, which means the vendor-neutral path is no longer a compromise path. It is the portability hedge that also interoperates with the platforms built on it.

A Minimal Span Sketch

The sketch below is illustrative and framework-free: a span emitter that records structure, timing, and status, and exports metadata only. Production systems should use an OpenTelemetry SDK or a platform SDK rather than hand-rolling, but the shape is the same everywhere.

import time
import uuid
from dataclasses import dataclass, field
from typing import Callable, Optional

RUN_ID = str(uuid.uuid4())

@dataclass
class Span:
    name: str
    run_id: str
    parent_id: Optional[str] = None
    started_at: float = 0.0
    attributes: dict[str, str] = field(default_factory=dict)
    status: str = "ok"

class SpanEmitter:
    """Illustrative. Production: an OTel SDK or platform SDK."""

    def __init__(self, export: Callable[[Span], None]) -> None:
        self._export = export
        self._stack: list[str] = []

    def start(self, name: str, **attributes: str) -> Span:
        span = Span(name=name, run_id=RUN_ID,
                    parent_id=self._stack[-1] if self._stack else None,
                    started_at=time.time(), attributes=attributes)
        self._stack.append(name)
        return span

    def end(self, span: Span, status: str = "ok") -> None:
        span.status = status
        self._stack.pop()
        self._export(span)

The attributes you pass are hashes, counters, and codes: tool="create_refund", args="sha256:...", attempt="2". Content capture, when a team opts into it for evaluation work, happens through a separate, redacted path rather than by loosening the default.

Debugging with Traces

Trace-first debugging is a fixed workflow, and its speed is its value:

  1. Find the run. By session, user hash, or failed-outcome filter. Ten seconds, because identity attributes exist on every span.
  2. Read the waterfall. The nesting shows the trajectory; the durations show where time went; the statuses show where it broke.
  3. Name the failing span. Wrong tool, failed call, timed-out retrieval, or a bad state transition. The span type maps to a subsystem.
  4. Fix the subsystem, not the symptom. A wrong-tool failure is a contract problem. A retry storm is a reliability problem. A context-overflow is a budgeting problem. The trace names the lane.

The same traces feed everything else: failed runs become scenarios for the evaluation suite, retry and failure spans drive the reliability work in agent reliability engineering, run-level cost and latency become the production health metrics of deployment architecture, and state-transition spans are the observable half of checkpointing and resumable workflows.

How Real Systems Do This

  • “Show me the trace” is the first sentence of every incident. Teams debug from the waterfall, not from log archaeology, and the run id is the incident’s primary key.
  • Content is off by default, everywhere. Metadata and codes on all spans; content capture is an explicit, redacted, per-workspace decision, usually limited to evaluation datasets.
  • Agent version is a first-class attribute. Dashboards slice success rate, cost, and latency by version, so regressions attribute to releases within minutes.
  • Alerts run on span aggregates. Success rate drops, cost per run spikes, retry storms, escalation rate changes, with thresholds tuned from the evaluation suite.
  • Vendor-neutral instrumentation under vendor platforms. Teams instrument with OpenTelemetry conventions and export to Langfuse or a backend of choice, keeping the option to change tools without re-instrumenting.

Decision Framework

  1. Who can host the telemetry? Compliance-sensitive estates self-host, which favors Langfuse or an OTel backend. Teams that want managed operations lean LangSmith.
  2. What is exported by default? Decide before the first trace: metadata, hashes, and codes. Content is opt-in.
  3. Which content is opt-in, and how is it redacted? Name the span types and the redaction point. If nobody can answer this, the default has already been chosen badly.
  4. What must the queries support? Find by run, session, user, outcome, and version. If a question matters in an incident, its attribute must exist on the span.
  5. What alerts exist? Success rate, cost, retries, escalation. Dashboards without alerts are postcards.
  6. How do traces feed evaluation? Failed runs should land in the scenario backlog with one action.
  7. What are retention and access rules? Written, enforced, and reviewed, because the tracing store is a sensitive database.

When NOT to Use This

  • Early prototypes. Before the loop stabilizes, structured logs cover the basics. Build tracing when the system starts surviving long enough to be debugged.
  • Ultra-low-latency hot paths. Instrument everything, sample aggressively, and never let the emitter sit on the critical path unsampled.
  • When the store cannot be secured. Tracing sensitive content into a system with weak access control is not observability. It is exfiltration with dashboards.
  • When nobody will read it. Traces nobody queries are storage cost. Start with the four dashboards your incidents actually use and grow from evidence of use.

Common Mistakes

  • Logging instead of tracing. The fallout: flat lines without causality, and incident reviews that reconstruct runs from timestamps by hand.
  • Exporting full content by default. The cost: credentials and personal data in a third-party backend, and a compliance incident with a very complete evidence trail.
  • No agent version on spans. What you get: a measurable regression that cannot be attributed to the release that caused it.
  • Traces without token and cost attributes. Where it lands: cost incidents invisible until the invoice, and no early warning anywhere.
  • No retention policy. The consequence: a telemetry archive that outlives its usefulness and outgrows its compliance story.
  • Traces that never feed evaluation. What follows: a free, continuous source of hard scenarios, ignored.

Key Takeaways

  • An agent run is a distributed workflow: one root span, nested spans for steps, model calls, tools, retrievals, state transitions, retries, failures, completion.
  • Record identity, timing, status, tokens, and outcome codes on every span. Agent version is non-negotiable for regression attribution.
  • Prompts, tool arguments, and results carry secrets and personal data. Content capture is opt-in, redacted, and never the default. Secure the store like a sensitive database.
  • LangSmith, Langfuse, and OpenTelemetry all express the same taxonomy. Ownership and integration decide the choice, and OTel conventions, including GenAI agent and MCP spans, make instrumentation portable.
  • Debug from the waterfall: find the run, read the trajectory, name the failing span, fix the subsystem it maps to.
  • Traces feed everything: evaluation scenarios, reliability metrics, deployment health, and audit evidence. They are the connective tissue of agent production.

FAQ

What is tracing for AI agents?

Recording an entire agent run as a hierarchy of spans: a root run with nested spans for steps, model calls, tool calls, retrievals, state transitions, retries, and completion, each carrying identity, timing, status, tokens, and outcome codes. It shows what the run did, in causal order, with costs attached.

Is OpenTelemetry suitable for LLM agents?

Yes. OpenTelemetry is a vendor-neutral standard for traces, metrics, and logs, and its semantic conventions now include generative AI coverage, with spans for agents and MCP. Instrumenting with OTel keeps your telemetry portable across backends, and platforms like Langfuse are built on it.

What is the difference between LangSmith and Langfuse?

LangSmith is a managed observability and evaluation platform with deep LangChain and LangGraph integration. Langfuse is an open-source AI engineering platform you can self-host, built on OpenTelemetry, with tracing, prompt management, evaluation, and annotation built in. Managed workflow versus self-hosted control is usually the deciding axis.

What should you never log in agent traces?

Raw prompts, tool arguments, tool results, and retrieved documents by default. They routinely contain secrets and personal data. Record hashes, counts, and codes instead, and make content capture an explicit, redacted opt-in for the span types that need it.

How is tracing different from logging for agents?

Logs are flat lines of text without structure or causality. A trace is a hierarchy with timing, parent-child relationships, and span attributes, which is what reconstructing a multi-step agent run requires. Logs can complement traces, but traces are the primary artifact for debugging runs.

Conclusion

Tracing is the cheapest discipline with the highest payoff in agent production. It is the difference between debugging a failed run in four minutes and reconstructing it from logs for four hours, and it is the raw material for evaluation, reliability engineering, cost control, and audit.

The stack matters less than the shape. Root run, steps, calls, transitions, retries, completion: nine span types, four attribute groups, and a privacy posture that records structure by default and content only by deliberate choice.

Trace everything the run does, record almost nothing it says.

Last updated on 3 October 2026

Share this article

Leave a Reply

Your email address will not be published. Required fields are marked *