Skip to content

Adversarial Review

Ratified

A reviewer whose job is to find holes, not to approve.

  • After generation, a separate agent reviews the output
  • Its job is to break the output, not to help it
  • It looks for flaws, gaps, unstated assumptions, missed edge cases, and claims without evidence
  • Give: the output, plus the spec or acceptance criteria
  • Withhold: the generator’s reasoning and chat history
  • Requirements let it judge correctness; withholding reasoning keeps it from inheriting the generator’s framing
  • LLM evaluators can score their own outputs higher than equal-quality outputs from others (Panickssery et al.)
  • LLM judges show self-enhancement, position, and verbosity biases (Zheng et al.)
  • Agents asked to grade their own work tend to praise it, even when quality is mediocre (Anthropic, harness design)
  • Tuning a standalone evaluator to be skeptical was more tractable than making the generator self-critical (same source)
  • Agree what “done” means per chunk before any code is written
  • Grade against a rubric with a hard threshold per criterion
  • Any criterion below threshold fails the chunk, with specific feedback
  • Exercise the running system where possible, not only the diff
  • Source for all four: the generator/evaluator harness in Anthropic, harness design
  • In that harness, on one app prompt, the full run took 6 hours; a single agent took 20 minutes (Anthropic, harness design)
  • The author reported the quality gap was immediately apparent; one prompt, not a benchmark
  • Decide by task value: review where being wrong costs more than the review
  • High-stakes output where a mistake is expensive
  • Output trusted without further human review
  • Plausible claims nobody has verified
  • Subjective quality that a rubric can make gradable
  • Low-stakes output
  • Cheap deterministic checks already cover the risk (Verification Loops)
  • Nobody will act on the findings

Prompt to the reviewer:

You are reviewing a pull request. Default to skepticism.
Inputs: the diff, and the ticket's acceptance criteria.
For each finding give: severity (high/medium/low), file:line,
the failure scenario (input -> wrong result), and a check
that would prove it. Do not report style preferences.
If you find nothing, say "no findings".

Output:

1. high api/refund.py:42 refund > original charge is accepted
verify: POST /refund amount=150 on a 100 charge -> expect 400
2. low api/refund.py:77 log line omits refund id
verify: grep log after refund -> id missing
  • Each finding is then verified: run the check
  • Finding 1 reproduces and blocks merge; finding 2 goes to the backlog
  • The same agent and context generating and reviewing
  • A soft brief: “check if this looks good”
  • The reviewer invents findings to look useful; require a verify step per finding
  • Findings never triaged, or ignored because they are inconvenient