Skip to content

Context Handoff

Ratified

Every session starts cold. Leave the next one a clean desk.

  • At a phase boundary, or before compaction, stop in a mergeable state
  • Commit the work
  • Write the handoff
  • The next session starts fresh and reads the handoff first
  • Goal: one line
  • Spec or plan: link
  • Done: each item with evidence (commit, test run)
  • In progress: what is partly built and where
  • Known broken: failing tests, open bugs
  • Decisions: what was chosen and why
  • Open questions: what needs a human
  • Next step: the single next action
  • Verify: commands to confirm the state
  • Compaction: summarizes earlier conversation in place; the same agent continues on a shorter history
  • Reset: clears the context and starts a new agent from a structured handoff
  • Compaction keeps continuity but gives no clean slate; a reset gives a clean slate but depends on the handoff holding enough state (Anthropic, harness design)
  • Compaction alone can leave the next session with a half-built, undocumented feature (Anthropic, long-running harnesses)
  • End of a plan phase
  • Context nearing its limit
  • Before switching model or agent
  • End of the working day
  • After two corrections on the same issue: start fresh with a better prompt (Session Rule; Claude Code best practices)
  • Read the handoff and recent commits
  • Pick the highest-priority unfinished item
  • Run a smoke test before new work, so inherited breakage is fixed first
  • Source: the same routine in Anthropic, long-running harnesses
  • State what must survive a summary
  • Example instruction: preserve the full list of modified files and any test commands (Claude Code best practices)
  • Models use information in the middle of long contexts less reliably than at the start or end (Lost in the Middle)
  • One harness report saw some models wrap up work early as they neared their perceived context limit (Anthropic, harness design); do not assume every model does this
  • Runs longer than one context window
  • Switching agent or model mid-run
  • A human takes over from an agent, or the reverse
  • Several agents take turns on one body of work
  • Tasks that fit in one session
  • Breadcrumbs on the board already hold the full state
  • Exploratory threads where the reasoning trail matters more than the state

Migrating a service from one ORM to another. Context reaches about 70%.

Goal: replace ORM in orders-service
Plan: docs/plans/orm-migration.md
Done: models ported (commit 3f1a), unit tests green (run #88)
In progress: repository layer, 4 of 9 files
Known broken: none
Decisions: keep soft-delete column; reporting job reads it
Open questions: none
Next: port OrderRepository.findByCustomer
Verify: make test && make migrate-check
  • The session resets; the new session reads the note, runs make test, then continues at findByCustomer
  • Without the note: a summary can keep “porting repositories” but drop “keep soft-delete column”
  • The next session could then remove the column and break reporting
  • Automatic summary as the only memory
  • Half-implemented feature with no note
  • A handoff without evidence links
  • Pasting the transcript as the handoff
  • Letting the agent rewrite acceptance tests in the progress file