Skip to content

Unattended Runs

Ratified

Decide when it stops before it starts.

  • Design up front, then a long stretch with no one watching
  • Human attention moves to written checkpoints and the end of the run
  • The run ends on a stop condition decided before launch, not on someone noticing
  • Spec: problem, constraints, acceptance criteria (Spec, Then Build)
  • Plan: milestones and progress on a work board or plan file the agent updates
  • Routing config: which model handles which work (Task Routing)
  • Time limit: a wall-clock cap on the run, enforced outside the agent (Cost Management)
  • Abort criteria: written, specific, countable
  • Stop condition: a time limit, or all acceptance criteria pass
  • Gates: irreversible actions blocked (Checkpoint Gates)
  • No babysitting
  • Read Progress Breadcrumbs at agreed times only
  • Checkpoints arrive as parked decisions, not live interruptions
  • Context resets hand off through the plan file (Context Handoff)
  • Checkable acceptance criteria exist
  • Verification is automated (Verification Loops)
  • Irreversible actions are gated
  • Exploratory work with no checkable win
  • No automated verification: the human becomes the loop
  • Irreversible actions with no gate
  • Task shorter than the setup

Setup

  • One ticket, 5-hour time limit

Run

  • Hour 0: spec and plan reviewed; abort rule “same integration test fails 3 times after distinct fixes”
  • Hours 0-3: milestones 1-3 pass; breadcrumbs updated after each
  • Hour 3: test_refund_idempotent fails a third time; abort fires

Abort report

  • Failed: test_refund_idempotent, duplicate refund on retry
  • Tried: request-ID dedupe in handler; DB unique constraint; transaction retry wrapper
  • Finding: payment client retries internally before the handler sees the request
  • Next step: human decides whether the client’s retry setting is in scope

Attention compared (illustrative)

  • Unattended: 20 minutes of spec review, 10 minutes reading the report
  • Chat-steered session on the same ticket: a check-in every few minutes for 3 hours
  • The report turns the next move into one decision
  • Longest unattended stretch: time between required human touches
  • Human touches per run: checkpoints, aborts, corrections
  • Tokens and attention per win: see Tokens to Value
  • Context: METR measures a 50% task-completion time horizon that doubled roughly every seven months from 2019 to 2025; its Time Horizon 1.1 update (Jan 2026) estimates a post-2023 doubling of about 131 days (165 under the original method). A longer horizon does not remove the need for stop rules
  • No stop condition
  • Uncapped retries
  • Checking in every ten minutes
  • Treating the timebox as a delivery deadline: the run stops at the cap, finished or not