Skip to content

Cost Management

Draft

The cost of a win is tokens plus the human minutes it took.

  • The unit is cost per accepted win, not monthly spend (TTV)
  • A win costs tokens plus human attention: steering, review, QA, rework
  • Monthly spend is a lagging view; it cannot tell a cheap win from a cheap failure
  • Human review and QA time count as cost, often the larger share (working judgment, to be measured)
  • One team building with agents found human QA capacity became the bottleneck, with human time and attention the fixed constraint (OpenAI)
  • Tag every model call with an outcome ID (ticket, PR, eval run)
  • Log human interventions and active minutes per outcome
  • Mark each outcome accepted as delivered, accepted after correction, or rejected; both accepted states count as wins and are reported apart (Evaluation)
  • Put tokens from failed or abandoned runs in a waste bucket, and drive it down
  • Report tokens per win and minutes per win side by side (Observability)
  • Retry cap: a fixed number of attempts before escalation
  • Time limit: a wall-clock cap on one agent run (Unattended Runs)
  • Breach behavior: stop, save state, and report; never retry silently past a limit (Fail Fast, Recover Smart)
  • Exempt critical paths explicitly, with a named owner
  • Send only what the step needs
  • Cache stable context such as system prompts and docs (Memory & Context)
  • Summarize long histories instead of resending them
  • Validate inputs before generation; reject bad inputs before spending tokens
  • Clear specs cut rework (Spec Then Build)
  • Cheap deterministic checks before expensive model checks (Verification Loops)
  • Parallel agents multiply tokens; justify fanout by task value (Subagent Fanout)
  • One vendor measured agents at about 4x the tokens of chat, and multi-agent systems at about 15x (Anthropic)
  • Group similar tasks to amortize setup overhead
  • Documented case: one prompt run solo and through a planner, generator, and evaluator harness (Anthropic)
    • Solo: 20 minutes; the core feature did not work
    • Full harness: 6 hours; the core feature worked
    • The cheaper run was not the cheaper win; a broken result is all waste
  • Illustrative comparison (hypothetical numbers, same task):
    • Fast-tier loop: 3 attempts, 40k tokens, 25 human minutes to review and fix
    • Capable tier: 1 attempt, 60k tokens, 5 human minutes to review
    • The capable run uses more tokens and costs less per win once minutes count
  • Use your own logs to fill in the real numbers before deciding
  • Optimizing cost without measuring value (TTV)
  • No visibility into spend breakdown
  • Hard limits that break critical paths
  • Ignoring cost until the bill arrives
  • Cutting tokens while human attention per win rises