Reproducibility
Define “similar” per use case. Never assume it.
The Principle
Section titled “The Principle”- Decide the reproducibility level each part of the system requires
- Build the mechanisms to achieve it: pinning, seeding, logging, evaluation
- Debugging: an unreproducible bug cannot be fixed or proven fixed
- Testing: assertions need a stable baseline, or tests flake and get deleted
- Auditing: “why did it do that?” requires replaying the inputs
- Trust: different answers to the same question erode belief in all of them
- Exact reproducibility everywhere is often unachievable and sometimes undesirable
Levels
Section titled “Levels”- Exact: byte-identical output; pinned model, temperature 0, fixed seed; breaks on provider updates
- Statistical: same distribution; assert on aggregates over N runs
- Behavioral: same kind of output (right answer, right tool call, valid schema); phrasing varies
Defaults
Section titled “Defaults”- Production: behavioral
- Debugging: exact, by replaying logged inputs and outputs, not regenerating
- Exploration and brainstorming: variance is a feature; mark those paths exempt
Signal of Violation
Section titled “Signal of Violation”- No level chosen: tests are retried until green, or their assertions get deleted
- Exact matching where behavioral fits: tests fail on rephrasing alone
- Debugging by regenerating: a bug is called fixed because a rerun happened to pass
- Model version and parameters not logged: behavior shifts and nobody can tell what changed
Implemented By
Section titled “Implemented By”- Prompt Regression Testing: behavioral baselines for prompts
- Structured Output: checkable output shape
- Observability & Logging: log model version, parameters, seed, prompt
- Evaluation & Benchmarking: N-run pass thresholds
Open Questions
Section titled “Open Questions”- What flake rate is acceptable before a suite loses trust?
- How do we detect a provider model update that shifts our baselines?
Proposal: Discussion #9
