Skip to content

Evaluation & Benchmarking

Draft

If you can’t measure it, you can’t improve it.

  • Detect regressions before users do
  • Compare alternatives objectively
  • Justify changes with data
  • Track TTV over time
  • Name what counts as success for each workload class: merged PR, resolved ticket, correct extraction
  • Write it down before building cases; every metric below divides by it
  • Success rate (did it work?)
  • Quality score (how good was it?)
  • Time to completion
  • Human intervention rate
  • Acceptance rate: share of delegated outputs accepted as delivered, counted apart from those accepted after correction
  • Corrections per task: correction turns, each tagged with its cause (Correction Diagnosis)
  • Diff size per accepted output: lines and files changed; large diffs tend to raise review cost
  • Tokens per outcome
  • Human attention per outcome (minutes, interventions) (TTV)
  • Retry rate
  • Context utilization
  • Latency (time to first token, total time)
  • Error rate
  • Availability
  • Offline evals: Run against a static dataset, compare outputs
  • Online evals: Monitor production traffic
  • A/B tests: Compare alternatives on live traffic
  • Human evals: Expert judgment on quality
  • Seed from real failures: every production miss or human correction becomes a case
  • Start small: about 20 cases drawn from real usage was enough to see large effects in one vendor’s early work (Anthropic)
  • Pass definition per case: exact, statistical, or behavioral (Reproducibility)
  • N-run thresholds: run each case several times; set the pass bar on the aggregate
  • Baseline first: record current scores before changing anything
  • Run on every change: prompt, skill, model, or tool change (Prompt Regression Testing)
  • Isolate trials: clean environment per run; shared state causes correlated failures (Anthropic)
  • pass@k: at least one of k attempts succeeds; rises with k
  • pass^k: all k attempts succeed; falls with k
  • 75% per-trial success gives about 42% pass^3 (Anthropic)
  • Use pass@k when one success is enough and a human or check picks it
  • Use pass^k when consistency is the product, such as a customer-facing agent
  • Grade the outcome in the environment (tests pass, record written), not the path taken
  • Checking exact tool-call sequences is brittle; agents find valid paths designers did not anticipate (Anthropic)
  • Keep the full transcript for every trial
  • Read transcripts regularly; a failure may be a grader bug, not an agent bug
  • Acceptable for open-ended output where no deterministic check exists
  • Prefer code-based graders when a deterministic check is possible
  • Calibrate against human labels before trusting scores; recheck periodically
  • Known biases: position, verbosity, and self-enhancement (Zheng et al.)
  • Same study: strong judges reached over 80% agreement with human preferences, about human-to-human level
  • Give the judge a rubric and an “unknown” option to reduce guessing (Anthropic)
  • Rerun the full suite on the candidate model before switching
  • Remove harness components one at a time to find which are still load-bearing (Anthropic)
  • Tie reruns to Deliberate Currency tripwires
  • Measure time and outcomes directly; self-reported speedup is unreliable
  • In one RCT, experienced open-source developers took 19% longer with AI tools, yet estimated a 20% speedup afterward (METR)
  • Limits: 16 developers on their own mature repositories, early-2025 tools; METR marks these results as out of date (kept here as history)
  • Follow-up (Feb 2026): METR judged its late-2025 rerun an unreliable signal, mainly because developers declined to work without AI; it believes speedup is likely higher now, but its data is only weak evidence of how much (METR)
  • Takeaway: measure your own workflows; published numbers age fast

A bounded way for one engineer to see whether delegation pays off, instead of estimating it.

  • Log per task: task type; mode (delegated, paired, by hand); total minutes, for every mode including by hand; for delegated and paired, minutes briefing and minutes reviewing and fixing; acceptance rate, corrections per task, post-merge fixes, diff size
  • Mix: at least 3 by-hand tasks per task type, alternating with delegated ones
  • Compare like tasks across modes, not across task types
  • Verdict: per task type, lower median total minutes wins; delegated wins only if post-merge fixes are no worse than by hand
  • Time bound: two weeks
  • Stop early on cause tags when one owns most corrections; keep logging minutes to the time bound
  • Abort if logging takes more than a few minutes per task; simplify the log, then restart
  • Output: the verdict per task type; the top cause tag is the next input to fix; share the log as evidence for the workload’s win definition (TTV)

The same protocol, run by a team on shared workflows.

  • Pick workflows: two or three recurring ones the team does weekly, such as bug fixes or endpoint additions
  • Name an owner: one person collects the logs and writes the report
  • Log: each engineer uses the Individual Baseline fields, plus workflow name
  • Mix: at least 3 by-hand tasks per workflow, alternating with delegated ones
  • Time bound: two weeks
  • Abort if fewer than half the team is logging after the first week, or logging takes more than a few minutes per task; simplify, then restart
  • Report fields: per workflow and mode, median total minutes, and median minutes briefing and reviewing; acceptance rate (delegated); corrections per task; post-merge fixes; top cause tag; one recommended input fix
  • Verdict: per workflow, by the Individual Baseline rule
  • Output: the report goes to the team’s discussion as evidence; the top fix gets an owner
  • Illustrative case (hypothetical numbers)
  • Task: extract priority, component, and repro steps from support tickets into a schema (Structured Output)
  • Cases: 20 real tickets, including 6 that previously produced wrong output
  • Pass per case: schema valid, priority and component exact match, repro steps judged present (behavioral)
  • Threshold: each case runs 5 times; suite passes at 95% of runs, no case below 4 of 5
  • Baseline: current model passes 98 of 100 runs
  • Model swap: cheaper candidate passes 88 of 100; three multi-issue tickets fail on component
  • Decision: swap blocked; the three tickets are added to the regression set
  • Vanity metrics (measuring what’s easy, not what matters)
  • No baseline (can’t tell if you’re improving)
  • Evaluating once (things drift)
  • Over-fitting to evals (gaming the metric)
  • Trusting a score without reading transcripts