Skip to content

Observability & Logging

Draft

You can’t fix what you can’t see.

  • Debug failures after the fact
  • Understand what the model actually did
  • Trace costs to specific operations
  • Detect drift and degradation
  • Input prompt (or hash if sensitive)
  • Output response
  • Model used, with exact model version
  • Parameters (temperature, reasoning effort, max tokens) and seed (Reproducibility)
  • Prompt or skill version
  • Outcome ID linking the call to a ticket, PR, or eval run (TTV)
  • Token counts (input, output, cached)
  • Latency
  • Success/failure status and finish reason
  • Conversation flow
  • Tool calls and results
  • Context accumulation
  • Total cost
  • Human interventions and active minutes
  • Retries and their triggers
  • Stop reason (completed, time limit hit, error, human stop)
  • Success rates over time
  • Token spend by task type
  • Tokens per win and human attention per win (Cost Management)
  • Waste bucket: tokens from failed or abandoned runs
  • Error patterns
  • Latency distributions
  • Trace tree: one root span per run; child spans per model call and per tool call
  • On each span: token counts, model, finish reason, tool name, and status
  • Standard attributes: target the vendor-neutral OpenTelemetry GenAI semantic conventions, which define agent, model-call, and tool-execution spans
  • The conventions are still in development; pin the version you emit
  • Human-readable trail: keep a short progress log separate from traces, for people resuming the work (Progress Breadcrumbs)
  • Read transcripts regularly: they show whether a failure is a real agent mistake or a harness or grader problem (Anthropic)
  • One team tuned its evaluator by reading its logs and fixing where its judgment diverged from a human’s (Anthropic)
  • Expose logs, metrics, and traces to the agent, not only to dashboards
  • An agent that can query its own run’s telemetry can verify and debug its own change
  • One team gave each worktree an ephemeral observability stack the agent queries directly, making goals like “startup under 800ms” checkable (OpenAI)
  • Scope agent access to the telemetry of its own run and environment
  • PII in prompts/responses
  • Retention policies
  • Access controls
  • Anonymization strategies
  • Logging everything (cost, privacy, noise)
  • Logging nothing (flying blind)
  • No correlation IDs (can’t trace across services)
  • Ignoring the logs you have
  • Logging tokens with no outcome ID (cost you cannot tie to a win)
  • Traces nobody reads