Model Selection & Routing
Pick models on your evals and cost per win, not on habit or headlines.
What This Covers
Section titled “What This Covers”- Model roster: the small set of models your org approves, grouped into tiers
- Routing policy: which tier handles which work, when to escalate, what to fall back to
- Gateway: the shared layer that enforces the policy across harnesses
- Model names change every few months; the slot and its rules should not
Key Principles
Section titled “Key Principles”- Use the smallest model that reliably succeeds at the task
- Route by task complexity, not by habit
- Re-evaluate routing decisions as model capabilities evolve
Selection Criteria
Section titled “Selection Criteria”- Capability on your evals: pass rate on your own cases (Evaluation), not a public leaderboard
- Cost per win, not per call: tokens plus human attention per accepted outcome (TTV)
- Latency: time to first token and total time, for interactive paths
- Context length: usable length on your inputs, not the advertised maximum
- Tool-use reliability: correct tool choice and valid arguments across many turns
- Reasoning effort dial: whether effort is adjustable per request, and what each level costs
- Data handling: retention, training use, region, and compliance terms
- Availability and fallback: rate limits, outage history, and a second source
| Tier | Typical work | Trade |
|---|---|---|
| Fast | Classification, extraction, formatting, search and read steps | Lowest cost and latency; fails on ambiguity |
| Balanced | Standard implementation, summarization, most tasks | Default for day-to-day work |
| Capable | Planning, architecture, review, security, hard debugging | Highest cost per call; fewest retries on hard work |
Routing Policy
Section titled “Routing Policy”- Default tier: one named default per workload class; deviations need a reason
- Escalation triggers: failed verification, repeated retry, low confidence, outward-facing or security-sensitive step (Step-Level Routing)
- Fallback chain: an ordered list per tier; on outage or rate limit, fail over or stop cleanly, never loop (Fail Fast, Recover Smart)
- Policy owner: a named owner sets defaults; teams override within bounds (Authority Cascade)
Gateway Responsibilities
Section titled “Gateway Responsibilities”- Central policy: one place to change defaults and fallbacks for every harness
- Limits: retry caps and time limits enforced at the gateway (Cost Management)
- Fallbacks: automatic failover along the configured chain
- Logging: model, version, and reasoning effort recorded on every call (Observability)
Re-evaluation
Section titled “Re-evaluation”- Trigger on Deliberate Currency tripwires, not on every release
- Rerun the eval suite on the candidate before switching
- Switch only when cost per win improves or holds at higher pass rate
- Recheck harness scaffolding too; a stronger model may make parts of it unnecessary (example)
Worked Example
Section titled “Worked Example”- Workload: ticket triage agent (read ticket, search code, draft a fix plan)
- Policy: fast tier for search and read steps; balanced tier drafts the plan; capable tier on a failed check
- Fallback: balanced tier falls back to a second provider’s balanced model, then stops with a clear error
- Gateway log: model and effort per step, tagged with the ticket ID
- Result to check: pass rate and minutes of human review per triaged ticket, before and after
Anti-patterns
Section titled “Anti-patterns”- Defaulting to the biggest model by habit
- One model for every step of a long trajectory
- Choosing on per-call price or public leaderboards alone
- No fallback when the preferred model is unavailable
- Calling a floating model alias the provider can update, then chasing a regression nobody logged
Related
Section titled “Related”- Task Routing: route per task
- Step-Level Routing: route per step inside a trajectory
- Prompt Regression Testing: prove a model switch is safe
- Cost Management: cost per accepted win
- Chain of Thought: when to raise reasoning effort
- Harness Selection: where model choice is exposed
Current Stack
Section titled “Current Stack”- Record the current model roster in the Current Stack Roster
