Skip to content

Model Selection & Routing

Draft

Pick models on your evals and cost per win, not on habit or headlines.

  • Model roster: the small set of models your org approves, grouped into tiers
  • Routing policy: which tier handles which work, when to escalate, what to fall back to
  • Gateway: the shared layer that enforces the policy across harnesses
  • Model names change every few months; the slot and its rules should not
  • Use the smallest model that reliably succeeds at the task
  • Route by task complexity, not by habit
  • Re-evaluate routing decisions as model capabilities evolve
  • Capability on your evals: pass rate on your own cases (Evaluation), not a public leaderboard
  • Cost per win, not per call: tokens plus human attention per accepted outcome (TTV)
  • Latency: time to first token and total time, for interactive paths
  • Context length: usable length on your inputs, not the advertised maximum
  • Tool-use reliability: correct tool choice and valid arguments across many turns
  • Reasoning effort dial: whether effort is adjustable per request, and what each level costs
  • Data handling: retention, training use, region, and compliance terms
  • Availability and fallback: rate limits, outage history, and a second source
Tier Typical work Trade
Fast Classification, extraction, formatting, search and read steps Lowest cost and latency; fails on ambiguity
Balanced Standard implementation, summarization, most tasks Default for day-to-day work
Capable Planning, architecture, review, security, hard debugging Highest cost per call; fewest retries on hard work
  • Default tier: one named default per workload class; deviations need a reason
  • Escalation triggers: failed verification, repeated retry, low confidence, outward-facing or security-sensitive step (Step-Level Routing)
  • Fallback chain: an ordered list per tier; on outage or rate limit, fail over or stop cleanly, never loop (Fail Fast, Recover Smart)
  • Policy owner: a named owner sets defaults; teams override within bounds (Authority Cascade)
  • Central policy: one place to change defaults and fallbacks for every harness
  • Limits: retry caps and time limits enforced at the gateway (Cost Management)
  • Fallbacks: automatic failover along the configured chain
  • Logging: model, version, and reasoning effort recorded on every call (Observability)
  • Trigger on Deliberate Currency tripwires, not on every release
  • Rerun the eval suite on the candidate before switching
  • Switch only when cost per win improves or holds at higher pass rate
  • Recheck harness scaffolding too; a stronger model may make parts of it unnecessary (example)
  • Workload: ticket triage agent (read ticket, search code, draft a fix plan)
  • Policy: fast tier for search and read steps; balanced tier drafts the plan; capable tier on a failed check
  • Fallback: balanced tier falls back to a second provider’s balanced model, then stops with a clear error
  • Gateway log: model and effort per step, tagged with the ticket ID
  • Result to check: pass rate and minutes of human review per triaged ticket, before and after
  • Defaulting to the biggest model by habit
  • One model for every step of a long trajectory
  • Choosing on per-call price or public leaderboards alone
  • No fallback when the preferred model is unavailable
  • Calling a floating model alias the provider can update, then chasing a regression nobody logged