Observability & Logging
You can’t fix what you can’t see.
Why Observability Matters
Section titled “Why Observability Matters”- Debug failures after the fact
- Understand what the model actually did
- Trace costs to specific operations
- Detect drift and degradation
What to Capture
Section titled “What to Capture”Per-Request
Section titled “Per-Request”- Input prompt (or hash if sensitive)
- Output response
- Model used, with exact model version
- Parameters (temperature, reasoning effort, max tokens) and seed (Reproducibility)
- Prompt or skill version
- Outcome ID linking the call to a ticket, PR, or eval run (TTV)
- Token counts (input, output, cached)
- Latency
- Success/failure status and finish reason
Per-Session
Section titled “Per-Session”- Conversation flow
- Tool calls and results
- Context accumulation
- Total cost
- Human interventions and active minutes
- Retries and their triggers
- Stop reason (completed, time limit hit, error, human stop)
Aggregate
Section titled “Aggregate”- Success rates over time
- Token spend by task type
- Tokens per win and human attention per win (Cost Management)
- Waste bucket: tokens from failed or abandoned runs
- Error patterns
- Latency distributions
Agent Runs
Section titled “Agent Runs”- Trace tree: one root span per run; child spans per model call and per tool call
- On each span: token counts, model, finish reason, tool name, and status
- Standard attributes: target the vendor-neutral OpenTelemetry GenAI semantic conventions, which define agent, model-call, and tool-execution spans
- The conventions are still in development; pin the version you emit
- Human-readable trail: keep a short progress log separate from traces, for people resuming the work (Progress Breadcrumbs)
- Read transcripts regularly: they show whether a failure is a real agent mistake or a harness or grader problem (Anthropic)
- One team tuned its evaluator by reading its logs and fixing where its judgment diverged from a human’s (Anthropic)
Legible to the Agent
Section titled “Legible to the Agent”- Expose logs, metrics, and traces to the agent, not only to dashboards
- An agent that can query its own run’s telemetry can verify and debug its own change
- One team gave each worktree an ephemeral observability stack the agent queries directly, making goals like “startup under 800ms” checkable (OpenAI)
- Scope agent access to the telemetry of its own run and environment
Current Stack
Section titled “Current Stack”- The org’s current observability picks live in the Current Stack Roster
Privacy Considerations
Section titled “Privacy Considerations”- PII in prompts/responses
- Retention policies
- Access controls
- Anonymization strategies
Anti-patterns
Section titled “Anti-patterns”- Logging everything (cost, privacy, noise)
- Logging nothing (flying blind)
- No correlation IDs (can’t trace across services)
- Ignoring the logs you have
- Logging tokens with no outcome ID (cost you cannot tie to a win)
- Traces nobody reads
Related
Section titled “Related”- Progress Breadcrumbs: the human-readable trail beside traces
- Cost Management: tokens and attention per win
- Evaluation & Benchmarking: transcripts show whether a failure is real
