Skip to content

Correction Diagnosis

Ratified

A correction made twice is an input bug.

There is no universal fix for correction-heavy output: the right approach depends on the task, model, codebase, available context, and how much ambiguity the agent must resolve alone, and the cause chain below is how to find which one applies.

  • The aim is a first result close to what you would have written; review is the backstop
  • When you correct output, find the input that produced the miss
  • Walk a fixed cause chain, upstream first, and stop at the first link that explains it
  • Fix that input, capture the fix where it loads next time, then rerun
  • You correct the same kind of output more than once
  • Correcting output takes longer than writing it would have
  • A session needs two corrections on one point
  • Exploration, where variance is the point (Reproducibility)
  • Throwaway output you will not reuse
  • Before walking the chain, classify the miss with the One-Off or Repeated rule
  • One-off: repair once and move on
  • Repeated: walk the chain below

Check top to bottom. An upstream cause usually makes downstream fixes useless.

Link Symptom Cause Fix
Task fit Faster to have written it by hand Wrong work delegated Delegate less or pair: Delegation Fit
Model fit Shallow reasoning on a hard step Tier too small for the step Raise the tier: Task Routing, Step-Level Routing
Context: missing Ignored a repo convention Convention not written down, or inconsistent Write it down: Consistency as Leverage, Standing Instructions
Context: wrong Output got worse after material was added Irrelevant or conflicting context Prune context: Memory & Context
Framing Solved the wrong problem No problem statement or constraints State the problem: Problem Before Prescription
Examples Wrong format or style No exemplar to mirror Add an exemplar: Few-Shot Examples, the Mirror field in Delegation Fit
Scope Diff too large to review Task too large for one brief Split the task: Agent Architecture, Delegation Fit
Execution Wandered mid-task No plan review or milestone Add a plan review: Spec, Then Build
Verification Plausible code that does not run No runnable check Add a runnable check: Verification Loops
Feedback Same fix made in two sessions Correction never captured Capture the fix: Discovery Propagation
  • Two corrections on the same point: stop correcting in chat
  • Rewrite the brief with what you learned, then start a fresh session
  • Failed attempts left in context steer the next attempt (Memory & Context)
  • Prior art: Claude Code best practices give the same rule: after two failed corrections, clear and write a better initial prompt
  • Tag each correction with its chain link
  • After a bounded period, the most frequent tag is the next fix
  • Protocol: Individual Baseline
  • Illustrative case (composite session, not a transcript)
  • Task: add a GET /invoices/{id}/pdf endpoint to a service with about 40 existing endpoints
  • Prompt: “Add an endpoint that returns the invoice as a PDF”
  • Corrections, each classified by Step 0:
    1. PDF footer shows the wrong page count; no repo convention involved (one-off: repaired once with the failing test output)
    2. Returned errors as plain strings; the repo wraps errors in apperr.New, and last week’s endpoint got the same correction (repeated: matches a known convention, second time; Context: missing)
  • Response: stop correcting; the convention exists only in reviewers’ heads
  • Input fix: one line proposed to the repo’s standing instructions, “Wrap errors with apperr.New”; once merged, it loads every session
  • Rerun in a fresh session, same prompt; until the proposal merges, the rerun’s brief carries the same line: no corrections
  • Minutes (illustrative; 40 by hand):
Path Minutes Against 40 by hand
First attempt, stopped at the repeated miss 2 prompt + 43 review and correction = 45 5 worse
Input fix 2 writing the line n/a
Rerun 2 prompt + 12 review = 14 26 better
Total on this task 45 + 2 + 14 = 61 21 worse
  • Verdict: delegating lost on this task
  • Payoff: the standing instruction loads next session, so the next similar endpoint costs about the rerun’s 14 minutes
  • Steering by chat correction instead of fixing the brief
  • Blaming the model before checking the inputs
  • Adding more context to fix a context problem
  • Never resetting a drifting session
  • Verifying first and framing never: a check catches the miss, but the wrong problem still gets solved