Skip to content

Verification Loops

Ratified

Trust but verify. Then fix what fails.

  • Generate output, then run a check
  • On failure, feed the specific failure back and try again
  • Stop when the check passes or a stop condition fires
  • Repair: the failure (error, failing test, diff) is new input to the next attempt
  • Retry: the same input sent again, hoping for a different result
  • A verification loop repairs; blind retry of bad output is what Fail Fast, Recover Smart rules out
  • Agents need ground truth from the environment, such as test results, at each step (Anthropic, building effective agents)
  • Deterministic: compile, tests, lint, typecheck, schema validation
  • End-to-end: drive the running system the way a user would
  • External reviewer: a separate agent or human grades against criteria (Adversarial Review)
  • Self-check: the generator judges its own work; weakest, since agents tend to praise their own output (Anthropic, harness design)
  • Each criterion names the command or test that proves it
  • Require evidence: the command and its output, not a claim of success
  • Keep the acceptance list as structured data; the agent may only flip a pass/fail field
  • The agent may not delete or edit tests; one harness used JSON because the model overwrote it less often than Markdown (Anthropic, long-running harnesses)
feedback = none
last_error = none
for attempt in 1..MAX:
output = generate(task, feedback)
result = run_checks(output) # deterministic first
log(attempt, result.summary)
if result.passed: return output, result.evidence
if result.error == last_error or time_limit_exceeded: break
feedback = result.error # specific, untruncated
last_error = result.error
escalate(failures, attempts, state)
  • Output has machine-checkable criteria
  • First-attempt accuracy is not reliable enough
  • A check costs less than the error it catches
  • No machine-checkable criteria exist
  • The check costs more than the error
  • The failure is a spec problem; fix the spec, not the output
  • Subjective quality without a rubric; use Adversarial Review

Task: fix failing test test_invoice_total_rounding. Cap: 3 attempts.

attempt 1: pytest tests/test_invoice.py -> FAIL expected 10.01, got 10.0
attempt 2: pytest tests/test_invoice.py -> FAIL expected 10.01, got 10.0
stop: same error twice
escalation: rounding happens in currency lib, not invoice code;
tried ROUND_HALF_UP in invoice.py twice; branch fix-rounding at a9c1
  • The human sees the cause in one read and decides whether to patch the library call site
  • No fourth attempt burns tokens on the same wrong file
  • The loop edits the test until it passes
  • Error output truncated, so the model repairs the wrong thing
  • The prompt grows every iteration; send the latest failure, not the history