Skip to content

Human in the Loop

Ratified

Automation with oversight, not automation as abandonment.

  • Write down which decisions need a human
  • Make every checkpoint fast to clear
  • Too little oversight ships irreversible mistakes (external email, destructive migration, compliance breach)
  • Too much oversight makes humans the bottleneck and review becomes rubber-stamping
  • The root failure is an undefined boundary; undefined lines drift toward convenience
  • High stakes: hard-to-reverse actions, external communication, significant spend, safety-critical changes
  • Low confidence: high uncertainty, novel situations, conflicting signals
  • Policy: compliance mandates, audit trails, formal approvals
  • Put the relevant context next to the decision
  • Present a recommended action, not an open question
  • Make approve/reject a single step
  • Batch similar low-risk decisions
  • For code and documents, ask for output shaped for review (Reviewable Output)
  • Undefined boundary: whether a human signs off depends on who ran the agent that day
  • Rubber-stamping: approval rate near 100%, with seconds spent per decision
  • Human as bottleneck: work waits in the approval queue longer than the agent took to do it
  • Oversight after the fact: an external message or destructive change is found with no recorded approval
  • Which signals show a checkpoint is mis-tuned (near-100% approval, queue latency)?
  • When no reviewer is available, does the system block, fall back, or queue?

Proposal: Discussion #8